topicscout 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- topicscout-0.1.0/LICENSE +21 -0
- topicscout-0.1.0/PKG-INFO +151 -0
- topicscout-0.1.0/README.md +138 -0
- topicscout-0.1.0/pyproject.toml +26 -0
- topicscout-0.1.0/setup.cfg +4 -0
- topicscout-0.1.0/tests/test_topicscout.py +198 -0
- topicscout-0.1.0/topicscout/__init__.py +2 -0
- topicscout-0.1.0/topicscout/__main__.py +3 -0
- topicscout-0.1.0/topicscout/cli.py +103 -0
- topicscout-0.1.0/topicscout/core.py +274 -0
- topicscout-0.1.0/topicscout/profiles/ai-memory.json +59 -0
- topicscout-0.1.0/topicscout/profiles/github-scrapers.json +55 -0
- topicscout-0.1.0/topicscout.egg-info/PKG-INFO +151 -0
- topicscout-0.1.0/topicscout.egg-info/SOURCES.txt +15 -0
- topicscout-0.1.0/topicscout.egg-info/dependency_links.txt +1 -0
- topicscout-0.1.0/topicscout.egg-info/entry_points.txt +2 -0
- topicscout-0.1.0/topicscout.egg-info/top_level.txt +1 -0
topicscout-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Adecubed
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: topicscout
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional.
|
|
5
|
+
Author: Adecubed
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/adecubed/topicscout
|
|
8
|
+
Keywords: github,scraper,search,discovery,topics
|
|
9
|
+
Requires-Python: >=3.11
|
|
10
|
+
Description-Content-Type: text/markdown
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Dynamic: license-file
|
|
13
|
+
|
|
14
|
+
# topicscout
|
|
15
|
+
|
|
16
|
+
[](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
|
|
17
|
+
[](https://pypi.org/project/topicscout/)
|
|
18
|
+
|
|
19
|
+
Find the GitHub repositories on a topic, and next time only the new ones.
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
|
|
23
|
+
topicscout run github-scrapers
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
`new.md` is a table of what appeared since the last run, most stars first: stars,
|
|
31
|
+
last push, language, license, a guess at the interface from the README (MCP, pip,
|
|
32
|
+
npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
|
|
33
|
+
so a weekly run is a short list, not the same 300 repos again.
|
|
34
|
+
|
|
35
|
+
No dependencies, Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
|
|
36
|
+
are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
|
|
37
|
+
|
|
38
|
+
## Why
|
|
39
|
+
|
|
40
|
+
GitHub search is good at "repos with this topic" and bad at "what's new since I last
|
|
41
|
+
looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
|
|
42
|
+
only tagged themselves". topicscout runs a fixed set of searches, applies your
|
|
43
|
+
filters, remembers what it has seen and what you already have.
|
|
44
|
+
|
|
45
|
+
It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
|
|
46
|
+
benchmark for AI memory systems, which uses it to find new memories to measure.
|
|
47
|
+
|
|
48
|
+
## Your own topic
|
|
49
|
+
|
|
50
|
+
Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
|
|
51
|
+
change the queries and the words a hit must contain, and run it:
|
|
52
|
+
|
|
53
|
+
```bash
|
|
54
|
+
topicscout run my-topic.json
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Run it again next week: `new.md` lists only what appeared in between. Put the repos
|
|
58
|
+
you already know or rejected in `scout-my-topic/known.txt`.
|
|
59
|
+
|
|
60
|
+
## From a script or an agent
|
|
61
|
+
|
|
62
|
+
`--json` prints the result on stdout, progress stays on stderr:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
topicscout run github-scrapers --json -q
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
```json
|
|
69
|
+
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
It is a command-line tool and a Python library (`from topicscout.core import Profile, run`),
|
|
73
|
+
not an MCP server; an agent calls it like any other command.
|
|
74
|
+
|
|
75
|
+
## Commands
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
|
|
79
|
+
[--known-from GLOB ...] [--json] [-q]
|
|
80
|
+
topicscout profiles # built-in profiles
|
|
81
|
+
topicscout doctor # token and remaining GitHub quota
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
- `PROFILE` is a built-in name or a path to your own JSON profile.
|
|
85
|
+
- `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
|
|
86
|
+
`new.md`, and your `known.txt`.
|
|
87
|
+
- `--json` prints the summary and the new repos as JSON on stdout; progress and
|
|
88
|
+
errors always go to stderr, so scripts and agents can read stdout as is.
|
|
89
|
+
|
|
90
|
+
## What counts as known
|
|
91
|
+
|
|
92
|
+
Repos you already have are marked `known` and never listed as new:
|
|
93
|
+
|
|
94
|
+
- `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
|
|
95
|
+
rejected repos here too.
|
|
96
|
+
- `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
|
|
97
|
+
files. adebench points it at its adapters, so a memory with an adapter drops out.
|
|
98
|
+
|
|
99
|
+
## Profiles
|
|
100
|
+
|
|
101
|
+
Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
|
|
102
|
+
(tools that scrape, crawl or mine GitHub). A profile is a JSON file:
|
|
103
|
+
|
|
104
|
+
```json
|
|
105
|
+
{
|
|
106
|
+
"name": "ai-memory",
|
|
107
|
+
"description": "Memory systems for AI agents and LLMs",
|
|
108
|
+
"queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
|
|
109
|
+
"exclude": "(?i)awesome|curated list|\\bpapers\\b",
|
|
110
|
+
"require": [
|
|
111
|
+
{"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
|
|
112
|
+
{"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
|
|
113
|
+
],
|
|
114
|
+
"interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
|
|
115
|
+
"min_stars": 10,
|
|
116
|
+
"days": 90
|
|
117
|
+
}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
- `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
|
|
121
|
+
queries. Each one always gets `fork:false archived:false stars:>=min_stars
|
|
122
|
+
pushed:>=today-days`; up to 300 results per query.
|
|
123
|
+
- `exclude`: a regex on name, description and topics; a match drops the repo.
|
|
124
|
+
- `require`: every pattern must match. With `own_words` it must be in the name or
|
|
125
|
+
description, not only in a topic: a database that tagged itself `agent-memory`
|
|
126
|
+
is not a memory. Hyphens and underscores count as spaces.
|
|
127
|
+
- `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
|
|
128
|
+
Leave it out and no README is fetched.
|
|
129
|
+
- `require_language` (default true) drops repos with no language, which are
|
|
130
|
+
usually lists and docs.
|
|
131
|
+
|
|
132
|
+
Patterns are Python regexes; add `(?i)` for case-insensitive.
|
|
133
|
+
|
|
134
|
+
## Rate limits
|
|
135
|
+
|
|
136
|
+
On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
|
|
137
|
+
stops, saves what it found and says so in `new.md` and on stderr. The next run
|
|
138
|
+
picks up the READMEs it did not get to.
|
|
139
|
+
|
|
140
|
+
## How it compares
|
|
141
|
+
|
|
142
|
+
- [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
|
|
143
|
+
(issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
|
|
144
|
+
topic.
|
|
145
|
+
- [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
|
|
146
|
+
lists the top repos for a keyword, once; topicscout adds filters, profiles and
|
|
147
|
+
memory across runs.
|
|
148
|
+
|
|
149
|
+
## License
|
|
150
|
+
|
|
151
|
+
MIT
|
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
# topicscout
|
|
2
|
+
|
|
3
|
+
[](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
|
|
4
|
+
[](https://pypi.org/project/topicscout/)
|
|
5
|
+
|
|
6
|
+
Find the GitHub repositories on a topic, and next time only the new ones.
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
|
|
10
|
+
topicscout run github-scrapers
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
```
|
|
14
|
+
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
`new.md` is a table of what appeared since the last run, most stars first: stars,
|
|
18
|
+
last push, language, license, a guess at the interface from the README (MCP, pip,
|
|
19
|
+
npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
|
|
20
|
+
so a weekly run is a short list, not the same 300 repos again.
|
|
21
|
+
|
|
22
|
+
No dependencies, Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
|
|
23
|
+
are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
|
|
24
|
+
|
|
25
|
+
## Why
|
|
26
|
+
|
|
27
|
+
GitHub search is good at "repos with this topic" and bad at "what's new since I last
|
|
28
|
+
looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
|
|
29
|
+
only tagged themselves". topicscout runs a fixed set of searches, applies your
|
|
30
|
+
filters, remembers what it has seen and what you already have.
|
|
31
|
+
|
|
32
|
+
It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
|
|
33
|
+
benchmark for AI memory systems, which uses it to find new memories to measure.
|
|
34
|
+
|
|
35
|
+
## Your own topic
|
|
36
|
+
|
|
37
|
+
Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
|
|
38
|
+
change the queries and the words a hit must contain, and run it:
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
topicscout run my-topic.json
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Run it again next week: `new.md` lists only what appeared in between. Put the repos
|
|
45
|
+
you already know or rejected in `scout-my-topic/known.txt`.
|
|
46
|
+
|
|
47
|
+
## From a script or an agent
|
|
48
|
+
|
|
49
|
+
`--json` prints the result on stdout, progress stays on stderr:
|
|
50
|
+
|
|
51
|
+
```bash
|
|
52
|
+
topicscout run github-scrapers --json -q
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
```json
|
|
56
|
+
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
It is a command-line tool and a Python library (`from topicscout.core import Profile, run`),
|
|
60
|
+
not an MCP server; an agent calls it like any other command.
|
|
61
|
+
|
|
62
|
+
## Commands
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
|
|
66
|
+
[--known-from GLOB ...] [--json] [-q]
|
|
67
|
+
topicscout profiles # built-in profiles
|
|
68
|
+
topicscout doctor # token and remaining GitHub quota
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
- `PROFILE` is a built-in name or a path to your own JSON profile.
|
|
72
|
+
- `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
|
|
73
|
+
`new.md`, and your `known.txt`.
|
|
74
|
+
- `--json` prints the summary and the new repos as JSON on stdout; progress and
|
|
75
|
+
errors always go to stderr, so scripts and agents can read stdout as is.
|
|
76
|
+
|
|
77
|
+
## What counts as known
|
|
78
|
+
|
|
79
|
+
Repos you already have are marked `known` and never listed as new:
|
|
80
|
+
|
|
81
|
+
- `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
|
|
82
|
+
rejected repos here too.
|
|
83
|
+
- `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
|
|
84
|
+
files. adebench points it at its adapters, so a memory with an adapter drops out.
|
|
85
|
+
|
|
86
|
+
## Profiles
|
|
87
|
+
|
|
88
|
+
Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
|
|
89
|
+
(tools that scrape, crawl or mine GitHub). A profile is a JSON file:
|
|
90
|
+
|
|
91
|
+
```json
|
|
92
|
+
{
|
|
93
|
+
"name": "ai-memory",
|
|
94
|
+
"description": "Memory systems for AI agents and LLMs",
|
|
95
|
+
"queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
|
|
96
|
+
"exclude": "(?i)awesome|curated list|\\bpapers\\b",
|
|
97
|
+
"require": [
|
|
98
|
+
{"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
|
|
99
|
+
{"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
|
|
100
|
+
],
|
|
101
|
+
"interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
|
|
102
|
+
"min_stars": 10,
|
|
103
|
+
"days": 90
|
|
104
|
+
}
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
- `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
|
|
108
|
+
queries. Each one always gets `fork:false archived:false stars:>=min_stars
|
|
109
|
+
pushed:>=today-days`; up to 300 results per query.
|
|
110
|
+
- `exclude`: a regex on name, description and topics; a match drops the repo.
|
|
111
|
+
- `require`: every pattern must match. With `own_words` it must be in the name or
|
|
112
|
+
description, not only in a topic: a database that tagged itself `agent-memory`
|
|
113
|
+
is not a memory. Hyphens and underscores count as spaces.
|
|
114
|
+
- `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
|
|
115
|
+
Leave it out and no README is fetched.
|
|
116
|
+
- `require_language` (default true) drops repos with no language, which are
|
|
117
|
+
usually lists and docs.
|
|
118
|
+
|
|
119
|
+
Patterns are Python regexes; add `(?i)` for case-insensitive.
|
|
120
|
+
|
|
121
|
+
## Rate limits
|
|
122
|
+
|
|
123
|
+
On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
|
|
124
|
+
stops, saves what it found and says so in `new.md` and on stderr. The next run
|
|
125
|
+
picks up the READMEs it did not get to.
|
|
126
|
+
|
|
127
|
+
## How it compares
|
|
128
|
+
|
|
129
|
+
- [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
|
|
130
|
+
(issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
|
|
131
|
+
topic.
|
|
132
|
+
- [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
|
|
133
|
+
lists the top repos for a keyword, once; topicscout adds filters, profiles and
|
|
134
|
+
memory across runs.
|
|
135
|
+
|
|
136
|
+
## License
|
|
137
|
+
|
|
138
|
+
MIT
|
|
@@ -0,0 +1,26 @@
|
|
|
1
|
+
[project]
|
|
2
|
+
name = "topicscout"
|
|
3
|
+
version = "0.1.0"
|
|
4
|
+
description = "Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional."
|
|
5
|
+
readme = "README.md"
|
|
6
|
+
requires-python = ">=3.11"
|
|
7
|
+
license = "MIT"
|
|
8
|
+
authors = [{name = "Adecubed"}]
|
|
9
|
+
dependencies = []
|
|
10
|
+
keywords = ["github", "scraper", "search", "discovery", "topics"]
|
|
11
|
+
|
|
12
|
+
[project.urls]
|
|
13
|
+
Homepage = "https://github.com/adecubed/topicscout"
|
|
14
|
+
|
|
15
|
+
[project.scripts]
|
|
16
|
+
topicscout = "topicscout.cli:main"
|
|
17
|
+
|
|
18
|
+
[build-system]
|
|
19
|
+
requires = ["setuptools>=77"]
|
|
20
|
+
build-backend = "setuptools.build_meta"
|
|
21
|
+
|
|
22
|
+
[tool.setuptools]
|
|
23
|
+
packages = ["topicscout", "topicscout.profiles"]
|
|
24
|
+
|
|
25
|
+
[tool.setuptools.package-data]
|
|
26
|
+
"topicscout.profiles" = ["*.json"]
|
|
@@ -0,0 +1,198 @@
|
|
|
1
|
+
"""topicscout with fake HTTP (no network)."""
|
|
2
|
+
from __future__ import annotations
|
|
3
|
+
|
|
4
|
+
import json
|
|
5
|
+
from datetime import date
|
|
6
|
+
from urllib.parse import parse_qs, urlparse
|
|
7
|
+
|
|
8
|
+
import pytest
|
|
9
|
+
|
|
10
|
+
from topicscout import cli, core
|
|
11
|
+
from topicscout.core import Profile
|
|
12
|
+
|
|
13
|
+
TODAY = date(2026, 9, 29)
|
|
14
|
+
MEM = Profile.load("ai-memory")
|
|
15
|
+
|
|
16
|
+
|
|
17
|
+
def repo(name, stars=50, pushed="2026-09-20T10:00:00Z", fork=False, archived=False,
|
|
18
|
+
language="Python", description="Long-term memory for agents", topics=("agent-memory",)):
|
|
19
|
+
return {"full_name": name, "html_url": f"https://github.com/{name}", "stargazers_count": stars,
|
|
20
|
+
"pushed_at": pushed, "fork": fork, "archived": archived, "language": language,
|
|
21
|
+
"description": description, "license": {"spdx_id": "MIT"}, "topics": list(topics)}
|
|
22
|
+
|
|
23
|
+
|
|
24
|
+
class FakeGitHub:
|
|
25
|
+
"""Every search query returns the same items; READMEs from a dict."""
|
|
26
|
+
|
|
27
|
+
def __init__(self, items, readmes=None, limited_after=None):
|
|
28
|
+
self.items, self.readmes = items, readmes or {}
|
|
29
|
+
self.limited_after = limited_after
|
|
30
|
+
self.searches, self.readme_calls = [], []
|
|
31
|
+
|
|
32
|
+
def __call__(self, url, accept):
|
|
33
|
+
path = urlparse(url).path
|
|
34
|
+
if path == "/search/repositories":
|
|
35
|
+
if self.limited_after is not None and len(self.searches) >= self.limited_after:
|
|
36
|
+
return 403, {"x-ratelimit-remaining": "0", "x-ratelimit-reset": "9999999999"}, "{}"
|
|
37
|
+
self.searches.append(parse_qs(urlparse(url).query)["q"][0])
|
|
38
|
+
return 200, {}, json.dumps({"items": self.items})
|
|
39
|
+
if path.endswith("/readme"):
|
|
40
|
+
name = path[len("/repos/"):-len("/readme")]
|
|
41
|
+
self.readme_calls.append(name)
|
|
42
|
+
if name in self.readmes:
|
|
43
|
+
return 200, {}, self.readmes[name]
|
|
44
|
+
return 404, {}, ""
|
|
45
|
+
raise AssertionError(url)
|
|
46
|
+
|
|
47
|
+
|
|
48
|
+
def run(tmp_path, gh, profile=MEM, **kw):
|
|
49
|
+
return core.run(profile, tmp_path / "scout", gh, today=TODAY, sleep=lambda s: None, **kw)
|
|
50
|
+
|
|
51
|
+
|
|
52
|
+
def load(tmp_path):
|
|
53
|
+
return json.loads((tmp_path / "scout" / "candidates.json").read_text(encoding="utf-8"))["repos"]
|
|
54
|
+
|
|
55
|
+
|
|
56
|
+
def one_query(**extra):
|
|
57
|
+
return Profile({"name": "t", "queries": ["topic:a"], **extra})
|
|
58
|
+
|
|
59
|
+
|
|
60
|
+
def test_queries_carry_the_filters():
|
|
61
|
+
gh = FakeGitHub([])
|
|
62
|
+
core.search(gh, ["topic:agent-memory"], min_stars=10, days=90, today=TODAY, sleep=lambda s: None)
|
|
63
|
+
q = gh.searches[0]
|
|
64
|
+
assert "topic:agent-memory" in q and "fork:false" in q and "archived:false" in q
|
|
65
|
+
assert "stars:>=10" in q and "pushed:>=2026-07-01" in q
|
|
66
|
+
|
|
67
|
+
|
|
68
|
+
def test_ai_memory_drops_forks_archived_stale_small_lists(tmp_path):
|
|
69
|
+
items = [repo("a/good"), repo("a/fork", fork=True), repo("a/old", archived=True),
|
|
70
|
+
repo("a/stale", pushed="2026-01-01T00:00:00Z"), repo("a/tiny", stars=3),
|
|
71
|
+
repo("a/nolang", language=None), repo("a/awesome-agent-memory"),
|
|
72
|
+
repo("a/papers", description="A curated list of memory papers"),
|
|
73
|
+
repo("a/db", description="Git for data, for agents", topics=("agent-memory", "database")),
|
|
74
|
+
repo("a/claude-mem", description="Persistent context across sessions for every agent")]
|
|
75
|
+
run(tmp_path, FakeGitHub(items))
|
|
76
|
+
assert set(load(tmp_path)) == {"a/good", "a/claude-mem"}
|
|
77
|
+
|
|
78
|
+
|
|
79
|
+
def test_github_scrapers_profile():
|
|
80
|
+
gs = Profile.load("github-scrapers")
|
|
81
|
+
cutoff = date(2025, 9, 29)
|
|
82
|
+
ok = repo("n/github-scraper", description="crawl GitHub web pages", topics=())
|
|
83
|
+
spam = repo("x/OnlySnap", description="Scrape OnlyFans, a GitHub mirror", topics=())
|
|
84
|
+
other = repo("y/web-scraper", description="scrape any website", topics=())
|
|
85
|
+
assert core.keep(ok, gs, min_stars=5, cutoff=cutoff)
|
|
86
|
+
assert not core.keep(spam, gs, min_stars=5, cutoff=cutoff)
|
|
87
|
+
assert not core.keep(other, gs, min_stars=5, cutoff=cutoff)
|
|
88
|
+
|
|
89
|
+
|
|
90
|
+
def test_require_own_words_ignores_topics():
|
|
91
|
+
p = one_query(require=[{"pattern": "(?i)memory", "own_words": True}])
|
|
92
|
+
tagged = repo("a/db", description="Git for data", topics=("agent-memory",))
|
|
93
|
+
said = repo("a/mem", description="memory for agents", topics=())
|
|
94
|
+
cutoff = date(2026, 1, 1)
|
|
95
|
+
assert not core.keep(tagged, p, min_stars=1, cutoff=cutoff)
|
|
96
|
+
assert core.keep(said, p, min_stars=1, cutoff=cutoff)
|
|
97
|
+
|
|
98
|
+
|
|
99
|
+
def test_require_language_can_be_off():
|
|
100
|
+
p = one_query(require_language=False)
|
|
101
|
+
assert core.keep(repo("a/docs", language=None), p, min_stars=1, cutoff=date(2026, 1, 1))
|
|
102
|
+
|
|
103
|
+
|
|
104
|
+
def test_interface_guess_from_readme():
|
|
105
|
+
g = lambda s: core.guess_interface(s, MEM) # noqa: E731
|
|
106
|
+
assert g("Add it to Claude: an MCP server (modelcontextprotocol)") == ["mcp"]
|
|
107
|
+
assert g("pip install foo\n\ncurl http://localhost:8080/v1/memories") == ["pip", "rest"]
|
|
108
|
+
assert g("npm install -g bar; docker compose up") == ["npm", "docker"]
|
|
109
|
+
assert g("A Claude Code plugin: /plugin install x") == ["plugin"]
|
|
110
|
+
assert g("") == []
|
|
111
|
+
|
|
112
|
+
|
|
113
|
+
def test_no_interfaces_no_readmes(tmp_path):
|
|
114
|
+
gh = FakeGitHub([repo("a/one")])
|
|
115
|
+
run(tmp_path, gh, profile=one_query())
|
|
116
|
+
assert gh.readme_calls == [] and load(tmp_path)["a/one"]["interface"] is None
|
|
117
|
+
|
|
118
|
+
|
|
119
|
+
def test_known_from_files_and_known_txt(tmp_path):
|
|
120
|
+
ad = tmp_path / "adapters"
|
|
121
|
+
ad.mkdir()
|
|
122
|
+
(ad / "x.py").write_text('"""Adapter for X (https://github.com/Owner/X-Mem)."""\n', encoding="utf-8")
|
|
123
|
+
(ad / "late.py").write_text("\n" * 30 + "# https://github.com/not/this\n", encoding="utf-8")
|
|
124
|
+
out = tmp_path / "scout"
|
|
125
|
+
out.mkdir()
|
|
126
|
+
(out / "known.txt").write_text("# rejected\nother/thing\n", encoding="utf-8")
|
|
127
|
+
known_from = [str(ad / "*.py")]
|
|
128
|
+
assert core.known_repos(out / "known.txt", known_from) == {"owner/x-mem", "other/thing"}
|
|
129
|
+
run(tmp_path, FakeGitHub([repo("owner/X-Mem"), repo("other/thing"), repo("new/one")]), known_from=known_from)
|
|
130
|
+
repos = load(tmp_path)
|
|
131
|
+
assert repos["owner/X-Mem"]["status"] == "known" and repos["new/one"]["status"] == "new"
|
|
132
|
+
md = (out / "new.md").read_text(encoding="utf-8")
|
|
133
|
+
assert "new/one" in md and "owner/X-Mem" not in md
|
|
134
|
+
|
|
135
|
+
|
|
136
|
+
def test_state_across_runs(tmp_path):
|
|
137
|
+
gh = FakeGitHub([repo("a/one", stars=20)], readmes={"a/one": "pip install one"})
|
|
138
|
+
first = run(tmp_path, gh)
|
|
139
|
+
assert [r["full_name"] for r in first["new"]] == ["a/one"] and gh.readme_calls == ["a/one"]
|
|
140
|
+
gh.items = [repo("a/one", stars=25), repo("b/two", stars=90)]
|
|
141
|
+
second = run(tmp_path, gh)
|
|
142
|
+
assert [r["full_name"] for r in second["new"]] == ["b/two"]
|
|
143
|
+
assert gh.readme_calls == ["a/one", "b/two"] # README only for first sightings
|
|
144
|
+
repos = load(tmp_path)
|
|
145
|
+
assert repos["a/one"]["status"] == "seen" and repos["a/one"]["stars"] == 25
|
|
146
|
+
assert repos["a/one"]["first_seen"] == "2026-09-29" and repos["a/one"]["interface"] == ["pip"]
|
|
147
|
+
|
|
148
|
+
|
|
149
|
+
def test_new_md_most_stars_first(tmp_path):
|
|
150
|
+
run(tmp_path, FakeGitHub([repo("a/small", stars=11), repo("a/big", stars=900)]))
|
|
151
|
+
md = (tmp_path / "scout" / "new.md").read_text(encoding="utf-8")
|
|
152
|
+
assert md.startswith("# New repos for ai-memory") and md.index("a/big") < md.index("a/small")
|
|
153
|
+
|
|
154
|
+
|
|
155
|
+
def test_readme_cap(tmp_path):
|
|
156
|
+
gh = FakeGitHub([repo(f"a/r{i}") for i in range(5)])
|
|
157
|
+
run(tmp_path, gh, readme_max=2)
|
|
158
|
+
assert len(gh.readme_calls) == 2
|
|
159
|
+
|
|
160
|
+
|
|
161
|
+
def test_rate_limit_saves_what_was_found(tmp_path):
|
|
162
|
+
gh = FakeGitHub([repo("a/one")], limited_after=1)
|
|
163
|
+
p = one_query()
|
|
164
|
+
p.queries = ["topic:a", "topic:b", "topic:c"]
|
|
165
|
+
summary = run(tmp_path, gh, profile=p)
|
|
166
|
+
assert summary["stopped"] and "rate limit" in summary["stopped"]
|
|
167
|
+
assert set(load(tmp_path)) == {"a/one"}
|
|
168
|
+
|
|
169
|
+
|
|
170
|
+
def test_profile_from_path_and_unknown(tmp_path):
|
|
171
|
+
f = tmp_path / "mine.json"
|
|
172
|
+
f.write_text(json.dumps({"queries": ["topic:x"]}), encoding="utf-8")
|
|
173
|
+
p = Profile.load(str(f))
|
|
174
|
+
assert p.name == "mine" and p.min_stars == 10 and p.days == 90
|
|
175
|
+
with pytest.raises(FileNotFoundError):
|
|
176
|
+
Profile.load("no-such-profile")
|
|
177
|
+
with pytest.raises(ValueError):
|
|
178
|
+
Profile({"queries": []})
|
|
179
|
+
|
|
180
|
+
|
|
181
|
+
def test_builtin_profiles_all_load():
|
|
182
|
+
names = core.builtin_profiles()
|
|
183
|
+
assert {"ai-memory", "github-scrapers"} <= set(names)
|
|
184
|
+
for n in names:
|
|
185
|
+
Profile.load(n)
|
|
186
|
+
|
|
187
|
+
|
|
188
|
+
def test_cli_run_json(tmp_path, monkeypatch, capsys):
|
|
189
|
+
monkeypatch.setattr(core, "fetch_github", FakeGitHub([repo("a/one")]))
|
|
190
|
+
monkeypatch.setattr(core.time, "sleep", lambda s: None)
|
|
191
|
+
rc = cli.main(["run", "ai-memory", "--out", str(tmp_path / "o"), "--json", "-q"])
|
|
192
|
+
out = json.loads(capsys.readouterr().out)
|
|
193
|
+
assert rc == 0 and out["profile"] == "ai-memory" and [r["full_name"] for r in out["new"]] == ["a/one"]
|
|
194
|
+
|
|
195
|
+
|
|
196
|
+
def test_cli_unknown_profile(capsys):
|
|
197
|
+
assert cli.main(["run", "nope"]) == 2
|
|
198
|
+
assert "no profile" in capsys.readouterr().err
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
"""topicscout run | profiles | doctor.
|
|
2
|
+
|
|
3
|
+
Results go to stdout (a summary line, or JSON with --json); progress and
|
|
4
|
+
errors go to stderr, so a script or an agent can read stdout as is.
|
|
5
|
+
"""
|
|
6
|
+
from __future__ import annotations
|
|
7
|
+
|
|
8
|
+
import argparse
|
|
9
|
+
import json
|
|
10
|
+
import os
|
|
11
|
+
import sys
|
|
12
|
+
from datetime import datetime, timezone
|
|
13
|
+
from pathlib import Path
|
|
14
|
+
|
|
15
|
+
from . import __version__
|
|
16
|
+
from .core import Profile, _stderr, builtin_profiles, rate_limit, run
|
|
17
|
+
|
|
18
|
+
|
|
19
|
+
def _run(args) -> int:
|
|
20
|
+
try:
|
|
21
|
+
profile = Profile.load(args.profile)
|
|
22
|
+
except (FileNotFoundError, ValueError, KeyError) as e:
|
|
23
|
+
_stderr(f"topicscout: {e}")
|
|
24
|
+
return 2
|
|
25
|
+
out = Path(args.out or f"scout-{profile.name}")
|
|
26
|
+
log = (lambda m: None) if args.quiet else _stderr
|
|
27
|
+
if not os.environ.get("GITHUB_TOKEN"):
|
|
28
|
+
log("no GITHUB_TOKEN: searches paced to 10 per minute, READMEs capped at 40")
|
|
29
|
+
s = run(profile, out, min_stars=args.min_stars, days=args.days, readme_max=args.readme_max,
|
|
30
|
+
known_from=args.known_from, log=log)
|
|
31
|
+
if args.json:
|
|
32
|
+
print(json.dumps({**s, "out": str(out)}, indent=2, ensure_ascii=False))
|
|
33
|
+
else:
|
|
34
|
+
print(f"{s['found']} found, {s['kept']} kept, {len(s['new'])} new -> {out / 'new.md'}")
|
|
35
|
+
if s["stopped"]:
|
|
36
|
+
_stderr(f"stopped early: {s['stopped']}")
|
|
37
|
+
return 0
|
|
38
|
+
|
|
39
|
+
|
|
40
|
+
def _profiles(args) -> int:
|
|
41
|
+
ps = builtin_profiles()
|
|
42
|
+
if args.json:
|
|
43
|
+
print(json.dumps(ps, indent=2))
|
|
44
|
+
else:
|
|
45
|
+
for name, desc in ps.items():
|
|
46
|
+
print(f"{name:20} {desc}")
|
|
47
|
+
return 0
|
|
48
|
+
|
|
49
|
+
|
|
50
|
+
def _doctor(args) -> int:
|
|
51
|
+
token = bool(os.environ.get("GITHUB_TOKEN"))
|
|
52
|
+
report = {"version": __version__, "token": token}
|
|
53
|
+
try:
|
|
54
|
+
res = rate_limit()
|
|
55
|
+
except Exception as e: # network, bad token: say it, don't crash
|
|
56
|
+
report["error"] = str(e)
|
|
57
|
+
else:
|
|
58
|
+
for key in ("search", "core"):
|
|
59
|
+
r = res[key]
|
|
60
|
+
reset = datetime.fromtimestamp(r["reset"], timezone.utc).strftime("%H:%M:%S UTC")
|
|
61
|
+
report[key] = {"remaining": r["remaining"], "limit": r["limit"], "reset": reset}
|
|
62
|
+
if args.json:
|
|
63
|
+
print(json.dumps(report, indent=2))
|
|
64
|
+
else:
|
|
65
|
+
print(f"topicscout {__version__}")
|
|
66
|
+
print(f"GITHUB_TOKEN: {'set' if token else 'not set (works, but slower and READMEs capped)'}")
|
|
67
|
+
if "error" in report:
|
|
68
|
+
print(f"GitHub: unreachable ({report['error']})")
|
|
69
|
+
else:
|
|
70
|
+
for key, label in (("search", "search API"), ("core", "READMEs (core API)")):
|
|
71
|
+
r = report[key]
|
|
72
|
+
print(f"{label}: {r['remaining']}/{r['limit']} left, resets {r['reset']}")
|
|
73
|
+
return 1 if "error" in report else 0
|
|
74
|
+
|
|
75
|
+
|
|
76
|
+
def main(argv: list[str] | None = None) -> int:
|
|
77
|
+
ap = argparse.ArgumentParser(prog="topicscout",
|
|
78
|
+
description="Find the GitHub repos on a topic, and only the new ones next time.")
|
|
79
|
+
ap.add_argument("--version", action="version", version=f"topicscout {__version__}")
|
|
80
|
+
sub = ap.add_subparsers(dest="cmd", required=True)
|
|
81
|
+
|
|
82
|
+
r = sub.add_parser("run", help="search GitHub with a profile and list what is new")
|
|
83
|
+
r.add_argument("profile", help="built-in profile name (see `topicscout profiles`) or path to a JSON profile")
|
|
84
|
+
r.add_argument("--out", help="state folder: candidates.json, new.md, known.txt (default scout-<profile>)")
|
|
85
|
+
r.add_argument("--min-stars", type=int, help="override the profile's min_stars")
|
|
86
|
+
r.add_argument("--days", type=int, help="override the profile's days: drop repos with no push in this many days")
|
|
87
|
+
r.add_argument("--readme-max", type=int, help="READMEs read per run (default 40, 300 with GITHUB_TOKEN)")
|
|
88
|
+
r.add_argument("--known-from", action="append", default=[], metavar="GLOB",
|
|
89
|
+
help="files whose first 20 lines name repos you already have (repeatable)")
|
|
90
|
+
r.add_argument("--json", action="store_true", help="print the summary and the new repos as JSON")
|
|
91
|
+
r.add_argument("-q", "--quiet", action="store_true", help="no progress on stderr")
|
|
92
|
+
r.set_defaults(func=_run)
|
|
93
|
+
|
|
94
|
+
p = sub.add_parser("profiles", help="list the built-in profiles")
|
|
95
|
+
p.add_argument("--json", action="store_true")
|
|
96
|
+
p.set_defaults(func=_profiles)
|
|
97
|
+
|
|
98
|
+
d = sub.add_parser("doctor", help="token and GitHub quota")
|
|
99
|
+
d.add_argument("--json", action="store_true")
|
|
100
|
+
d.set_defaults(func=_doctor)
|
|
101
|
+
|
|
102
|
+
args = ap.parse_args(argv)
|
|
103
|
+
return args.func(args)
|
|
@@ -0,0 +1,274 @@
|
|
|
1
|
+
"""Search, filter, remember: the whole scout, with no dependencies.
|
|
2
|
+
|
|
3
|
+
A profile says what to look for (GitHub search queries), what a real hit
|
|
4
|
+
must say about itself (require), what to drop (exclude) and how to guess
|
|
5
|
+
the interface from the README. The state is kept across runs, so new.md
|
|
6
|
+
holds only what appeared since the last one.
|
|
7
|
+
"""
|
|
8
|
+
from __future__ import annotations
|
|
9
|
+
|
|
10
|
+
import glob
|
|
11
|
+
import json
|
|
12
|
+
import os
|
|
13
|
+
import re
|
|
14
|
+
import sys
|
|
15
|
+
import time
|
|
16
|
+
import urllib.error
|
|
17
|
+
import urllib.request
|
|
18
|
+
from datetime import date, datetime, timedelta, timezone
|
|
19
|
+
from pathlib import Path
|
|
20
|
+
from urllib.parse import urlencode
|
|
21
|
+
|
|
22
|
+
API = "https://api.github.com"
|
|
23
|
+
GITHUB_URL = re.compile(r"github\.com/([\w.-]+)/([\w.-]+)")
|
|
24
|
+
|
|
25
|
+
|
|
26
|
+
class RateLimited(Exception):
|
|
27
|
+
pass
|
|
28
|
+
|
|
29
|
+
|
|
30
|
+
def fetch_github(url: str, accept: str) -> tuple[int, dict, str]:
|
|
31
|
+
"""GET with the optional GITHUB_TOKEN; returns (status, lower-case headers, body)."""
|
|
32
|
+
headers = {"Accept": accept, "User-Agent": "topicscout", "X-GitHub-Api-Version": "2022-11-28"}
|
|
33
|
+
if os.environ.get("GITHUB_TOKEN"):
|
|
34
|
+
headers["Authorization"] = f"Bearer {os.environ['GITHUB_TOKEN']}"
|
|
35
|
+
req = urllib.request.Request(url, headers=headers)
|
|
36
|
+
try:
|
|
37
|
+
with urllib.request.urlopen(req, timeout=30) as r:
|
|
38
|
+
return r.status, {k.lower(): v for k, v in r.headers.items()}, r.read().decode("utf-8", "replace")
|
|
39
|
+
except urllib.error.HTTPError as e:
|
|
40
|
+
return e.code, {k.lower(): v for k, v in e.headers.items()}, e.read().decode("utf-8", "replace")
|
|
41
|
+
|
|
42
|
+
|
|
43
|
+
def _stderr(msg: str) -> None:
|
|
44
|
+
print(msg, file=sys.stderr, flush=True)
|
|
45
|
+
|
|
46
|
+
|
|
47
|
+
# ---------------------------------------------------------------- profiles
|
|
48
|
+
|
|
49
|
+
PROFILES = Path(__file__).resolve().parent / "profiles"
|
|
50
|
+
|
|
51
|
+
|
|
52
|
+
def builtin_profiles() -> dict[str, str]:
|
|
53
|
+
"""name -> description of the profiles shipped with the package."""
|
|
54
|
+
out = {}
|
|
55
|
+
for p in sorted(PROFILES.glob("*.json")):
|
|
56
|
+
out[p.stem] = json.loads(p.read_text(encoding="utf-8")).get("description", "")
|
|
57
|
+
return out
|
|
58
|
+
|
|
59
|
+
|
|
60
|
+
class Profile:
|
|
61
|
+
"""A compiled profile: queries, require/exclude patterns, interface guesses."""
|
|
62
|
+
|
|
63
|
+
def __init__(self, data: dict):
|
|
64
|
+
self.name = data.get("name", "profile")
|
|
65
|
+
self.description = data.get("description", "")
|
|
66
|
+
self.queries = list(data["queries"])
|
|
67
|
+
if not self.queries:
|
|
68
|
+
raise ValueError("a profile needs at least one query")
|
|
69
|
+
self.exclude = re.compile(data["exclude"]) if data.get("exclude") else None
|
|
70
|
+
# own_words: the pattern must be in the repo's name or description, not only in a topic
|
|
71
|
+
self.require = [(re.compile(r["pattern"]), bool(r.get("own_words"))) for r in data.get("require", [])]
|
|
72
|
+
self.interfaces = [(n, re.compile(p)) for n, p in data.get("interfaces", [])]
|
|
73
|
+
self.min_stars = int(data.get("min_stars", 10))
|
|
74
|
+
self.days = int(data.get("days", 90))
|
|
75
|
+
self.require_language = bool(data.get("require_language", True))
|
|
76
|
+
|
|
77
|
+
@classmethod
|
|
78
|
+
def load(cls, name_or_path: str) -> "Profile":
|
|
79
|
+
path = Path(name_or_path)
|
|
80
|
+
if not path.is_file():
|
|
81
|
+
path = PROFILES / f"{name_or_path}.json"
|
|
82
|
+
if not path.is_file():
|
|
83
|
+
raise FileNotFoundError(f"no profile {name_or_path!r}: not a file, and built-in ones are "
|
|
84
|
+
f"{', '.join(builtin_profiles())}")
|
|
85
|
+
data = json.loads(path.read_text(encoding="utf-8"))
|
|
86
|
+
data.setdefault("name", path.stem)
|
|
87
|
+
return cls(data)
|
|
88
|
+
|
|
89
|
+
|
|
90
|
+
# ---------------------------------------------------------------- search
|
|
91
|
+
|
|
92
|
+
def _get(fetch, url, accept, sleep):
|
|
93
|
+
"""One request; waits out a short rate limit once, raises RateLimited on a long one."""
|
|
94
|
+
for attempt in (0, 1):
|
|
95
|
+
status, headers, body = fetch(url, accept)
|
|
96
|
+
limited = status == 429 or (status == 403 and (headers.get("x-ratelimit-remaining") == "0"
|
|
97
|
+
or "retry-after" in headers))
|
|
98
|
+
if not limited:
|
|
99
|
+
return status, body
|
|
100
|
+
if "retry-after" in headers:
|
|
101
|
+
wait = int(headers["retry-after"])
|
|
102
|
+
else:
|
|
103
|
+
wait = int(headers.get("x-ratelimit-reset", "0")) - int(time.time())
|
|
104
|
+
if attempt or wait > 70:
|
|
105
|
+
raise RateLimited(f"GitHub rate limit, reset in {max(wait, 0)} s")
|
|
106
|
+
sleep(max(wait, 0) + 1)
|
|
107
|
+
raise AssertionError("unreachable")
|
|
108
|
+
|
|
109
|
+
|
|
110
|
+
def search(fetch, queries, *, min_stars, days, today, sleep, pause=0.0, pages=3,
|
|
111
|
+
log=lambda m: None) -> tuple[dict, str | None]:
|
|
112
|
+
"""Run the queries; returns ({full_name: item}, reason it stopped early or None)."""
|
|
113
|
+
since = (today - timedelta(days=days)).isoformat()
|
|
114
|
+
found: dict[str, dict] = {}
|
|
115
|
+
first = True
|
|
116
|
+
for i, q in enumerate(queries, 1):
|
|
117
|
+
full = f"{q} fork:false archived:false stars:>={min_stars} pushed:>={since}"
|
|
118
|
+
for page in range(1, pages + 1):
|
|
119
|
+
if not first and pause:
|
|
120
|
+
sleep(pause)
|
|
121
|
+
first = False
|
|
122
|
+
url = f"{API}/search/repositories?" + urlencode(
|
|
123
|
+
{"q": full, "sort": "stars", "order": "desc", "per_page": 100, "page": page})
|
|
124
|
+
try:
|
|
125
|
+
status, body = _get(fetch, url, "application/vnd.github+json", sleep)
|
|
126
|
+
except RateLimited as e:
|
|
127
|
+
return found, str(e)
|
|
128
|
+
if status != 200:
|
|
129
|
+
return found, f"search failed ({status}) on {q!r}: {body[:200]}"
|
|
130
|
+
items = json.loads(body).get("items", [])
|
|
131
|
+
for it in items:
|
|
132
|
+
found.setdefault(it["full_name"], it)
|
|
133
|
+
log(f"[{i}/{len(queries)}] {q} page {page}: {len(items)} ({len(found)} so far)")
|
|
134
|
+
if len(items) < 100:
|
|
135
|
+
break
|
|
136
|
+
return found, None
|
|
137
|
+
|
|
138
|
+
|
|
139
|
+
def keep(item: dict, profile: Profile, *, min_stars: int, cutoff: date) -> bool:
|
|
140
|
+
if item.get("fork") or item.get("archived"):
|
|
141
|
+
return False
|
|
142
|
+
if profile.require_language and not item.get("language"):
|
|
143
|
+
return False
|
|
144
|
+
if item.get("stargazers_count", 0) < min_stars:
|
|
145
|
+
return False
|
|
146
|
+
if datetime.fromisoformat(item["pushed_at"].replace("Z", "+00:00")).date() < cutoff:
|
|
147
|
+
return False
|
|
148
|
+
own = f"{item['full_name']} {item.get('description') or ''}"
|
|
149
|
+
text = f"{own} {' '.join(item.get('topics') or [])}"
|
|
150
|
+
if profile.exclude and profile.exclude.search(text):
|
|
151
|
+
return False
|
|
152
|
+
own_n, text_n = (s.replace("-", " ").replace("_", " ") for s in (own, text))
|
|
153
|
+
return all(rx.search(own_n if own_words else text_n) for rx, own_words in profile.require)
|
|
154
|
+
|
|
155
|
+
|
|
156
|
+
def guess_interface(readme: str, profile: Profile) -> list[str]:
|
|
157
|
+
return [name for name, rx in profile.interfaces if rx.search(readme)]
|
|
158
|
+
|
|
159
|
+
|
|
160
|
+
def known_repos(known_file: Path, known_from: list[str] = (), head: int = 20) -> set[str]:
|
|
161
|
+
"""owner/repo (lower case) from known.txt plus the GitHub URLs in the first
|
|
162
|
+
`head` lines of the files matching the known_from globs."""
|
|
163
|
+
known: set[str] = set()
|
|
164
|
+
for pattern in known_from:
|
|
165
|
+
for f in glob.glob(pattern, recursive=True):
|
|
166
|
+
p = Path(f)
|
|
167
|
+
if not p.is_file():
|
|
168
|
+
continue
|
|
169
|
+
top = "\n".join(p.read_text(encoding="utf-8", errors="replace").splitlines()[:head])
|
|
170
|
+
for owner, name in GITHUB_URL.findall(top):
|
|
171
|
+
name = name.rstrip(".").removesuffix(".git")
|
|
172
|
+
known.add(f"{owner}/{name}".lower())
|
|
173
|
+
if known_file.is_file():
|
|
174
|
+
for line in known_file.read_text(encoding="utf-8").splitlines():
|
|
175
|
+
line = line.split("#", 1)[0].strip()
|
|
176
|
+
if line:
|
|
177
|
+
known.add(line.lower())
|
|
178
|
+
return known
|
|
179
|
+
|
|
180
|
+
|
|
181
|
+
# ---------------------------------------------------------------- run
|
|
182
|
+
|
|
183
|
+
def _record(item: dict) -> dict:
|
|
184
|
+
return {"full_name": item["full_name"], "url": item["html_url"], "stars": item["stargazers_count"],
|
|
185
|
+
"pushed": item["pushed_at"][:10], "license": (item.get("license") or {}).get("spdx_id"),
|
|
186
|
+
"language": item.get("language"), "description": item.get("description") or "",
|
|
187
|
+
"topics": item.get("topics") or []}
|
|
188
|
+
|
|
189
|
+
|
|
190
|
+
def _new_md(new: list[dict], profile: Profile, today: date, stopped: str | None) -> str:
|
|
191
|
+
lines = [f"# New repos for {profile.name}, {today.isoformat()}", ""]
|
|
192
|
+
if stopped:
|
|
193
|
+
lines += [f"> Stopped early: {stopped}. The list is partial.", ""]
|
|
194
|
+
if not new:
|
|
195
|
+
return "\n".join(lines + ["Nothing new since the last run.", ""])
|
|
196
|
+
lines += [f"{len(new)} new.", "", "| repo | stars | last push | language | license | interface | description |",
|
|
197
|
+
"|---|---:|---|---|---|---|---|"]
|
|
198
|
+
for r in new:
|
|
199
|
+
iface = ", ".join(r["interface"]) if r.get("interface") else ("?" if r.get("interface") is None else "-")
|
|
200
|
+
desc = r["description"].replace("|", "\\|").replace("\n", " ")[:160]
|
|
201
|
+
lines.append(f"| [{r['full_name']}]({r['url']}) | {r['stars']} | {r['pushed']} | {r['language'] or '-'} | "
|
|
202
|
+
f"{r['license'] or '-'} | {iface} | {desc} |")
|
|
203
|
+
return "\n".join(lines + [""])
|
|
204
|
+
|
|
205
|
+
|
|
206
|
+
def run(profile: Profile, out: Path, fetch=None, *, today: date | None = None, sleep=None,
|
|
207
|
+
min_stars: int | None = None, days: int | None = None, readme_max: int | None = None,
|
|
208
|
+
pause: float | None = None, known_from: list[str] = (), log=lambda m: None) -> dict:
|
|
209
|
+
fetch = fetch or fetch_github
|
|
210
|
+
sleep = sleep or time.sleep
|
|
211
|
+
today = today or datetime.now(timezone.utc).date()
|
|
212
|
+
min_stars = profile.min_stars if min_stars is None else min_stars
|
|
213
|
+
days = profile.days if days is None else days
|
|
214
|
+
token = bool(os.environ.get("GITHUB_TOKEN"))
|
|
215
|
+
if pause is None:
|
|
216
|
+
pause = 2.0 if token else 7.0 # search API: 30/min with a token, 10/min without
|
|
217
|
+
if readme_max is None:
|
|
218
|
+
readme_max = 300 if token else 40
|
|
219
|
+
out.mkdir(parents=True, exist_ok=True)
|
|
220
|
+
state_file = out / "candidates.json"
|
|
221
|
+
state = json.loads(state_file.read_text(encoding="utf-8")) if state_file.is_file() else {"repos": {}}
|
|
222
|
+
repos: dict[str, dict] = state["repos"]
|
|
223
|
+
known = known_repos(out / "known.txt", known_from)
|
|
224
|
+
|
|
225
|
+
found, stopped = search(fetch, profile.queries, min_stars=min_stars, days=days, today=today,
|
|
226
|
+
sleep=sleep, pause=pause, log=log)
|
|
227
|
+
cutoff = today - timedelta(days=days)
|
|
228
|
+
new_names = []
|
|
229
|
+
for name, item in found.items():
|
|
230
|
+
if not keep(item, profile, min_stars=min_stars, cutoff=cutoff):
|
|
231
|
+
continue
|
|
232
|
+
rec = _record(item)
|
|
233
|
+
old = repos.get(name)
|
|
234
|
+
rec["first_seen"] = old["first_seen"] if old else today.isoformat()
|
|
235
|
+
rec["last_seen"] = today.isoformat()
|
|
236
|
+
rec["interface"] = old.get("interface") if old else None
|
|
237
|
+
if name.lower() in known:
|
|
238
|
+
rec["status"] = "known"
|
|
239
|
+
elif old:
|
|
240
|
+
rec["status"] = "seen"
|
|
241
|
+
else:
|
|
242
|
+
rec["status"] = "new"
|
|
243
|
+
new_names.append(name)
|
|
244
|
+
repos[name] = rec
|
|
245
|
+
|
|
246
|
+
# READMEs: first sightings first, then older ones never read (cap reached last time)
|
|
247
|
+
if profile.interfaces:
|
|
248
|
+
todo = new_names + [n for n, r in repos.items() if r["status"] == "seen" and r.get("interface") is None]
|
|
249
|
+
for name in todo[:readme_max]:
|
|
250
|
+
try:
|
|
251
|
+
status, body = _get(fetch, f"{API}/repos/{name}/readme", "application/vnd.github.raw", sleep)
|
|
252
|
+
except RateLimited as e:
|
|
253
|
+
stopped = stopped or f"{e} (while reading READMEs)"
|
|
254
|
+
break
|
|
255
|
+
repos[name]["interface"] = guess_interface(body, profile) if status == 200 else []
|
|
256
|
+
log(f"READMEs read: {min(len(todo), readme_max)} of {len(todo)}")
|
|
257
|
+
|
|
258
|
+
state = {"profile": profile.name, "updated": today.isoformat(),
|
|
259
|
+
"repos": dict(sorted(repos.items(), key=lambda kv: kv[0].lower()))}
|
|
260
|
+
state_file.write_text(json.dumps(state, indent=2, ensure_ascii=False) + "\n", encoding="utf-8")
|
|
261
|
+
new = sorted((repos[n] for n in new_names), key=lambda r: -r["stars"])
|
|
262
|
+
(out / "new.md").write_text(_new_md(new, profile, today, stopped), encoding="utf-8")
|
|
263
|
+
return {"profile": profile.name, "found": len(found),
|
|
264
|
+
"kept": sum(1 for r in repos.values() if r["last_seen"] == today.isoformat()),
|
|
265
|
+
"new": new, "stopped": stopped}
|
|
266
|
+
|
|
267
|
+
|
|
268
|
+
def rate_limit(fetch=None) -> dict:
|
|
269
|
+
"""GitHub's own view of the quota; this call does not count against it."""
|
|
270
|
+
fetch = fetch or fetch_github
|
|
271
|
+
status, _, body = fetch(f"{API}/rate_limit", "application/vnd.github+json")
|
|
272
|
+
if status != 200:
|
|
273
|
+
raise RuntimeError(f"rate_limit failed ({status}): {body[:200]}")
|
|
274
|
+
return json.loads(body)["resources"]
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "ai-memory",
|
|
3
|
+
"description": "Memory systems for AI agents and LLMs (the adebench list)",
|
|
4
|
+
"queries": [
|
|
5
|
+
"topic:agent-memory",
|
|
6
|
+
"topic:ai-memory",
|
|
7
|
+
"topic:llm-memory",
|
|
8
|
+
"topic:memory-mcp",
|
|
9
|
+
"topic:mcp-memory",
|
|
10
|
+
"topic:long-term-memory",
|
|
11
|
+
"topic:agentic-memory",
|
|
12
|
+
"topic:memory-layer",
|
|
13
|
+
"topic:memory-system",
|
|
14
|
+
"\"agent memory\" in:name,description",
|
|
15
|
+
"\"llm memory\" in:name,description",
|
|
16
|
+
"\"memory layer\" in:name,description",
|
|
17
|
+
"\"mcp memory\" in:name,description",
|
|
18
|
+
"\"long-term memory\" in:description",
|
|
19
|
+
"\"persistent memory\" in:description"
|
|
20
|
+
],
|
|
21
|
+
"exclude": "(?i)awesome|curated list|reading list|\\bpapers\\b|\\bsurvey\\b",
|
|
22
|
+
"require": [
|
|
23
|
+
{
|
|
24
|
+
"pattern": "(?i)memor|\\bmem\\b|remember",
|
|
25
|
+
"own_words": true
|
|
26
|
+
},
|
|
27
|
+
{
|
|
28
|
+
"pattern": "(?i)\\b(agents?|agentic|llms?|ai|mcp|claude|gpt|rag|assistants?|chatbots?)\\b"
|
|
29
|
+
}
|
|
30
|
+
],
|
|
31
|
+
"interfaces": [
|
|
32
|
+
[
|
|
33
|
+
"mcp",
|
|
34
|
+
"(?i)\\bmcp\\b|modelcontextprotocol"
|
|
35
|
+
],
|
|
36
|
+
[
|
|
37
|
+
"pip",
|
|
38
|
+
"(?i)\\bpip3? install\\b|\\buv add\\b|\\bpypi\\b"
|
|
39
|
+
],
|
|
40
|
+
[
|
|
41
|
+
"npm",
|
|
42
|
+
"(?i)\\bnpm (i|install)\\b|\\bnpx\\b|\\bbun add\\b|\\bpnpm add\\b"
|
|
43
|
+
],
|
|
44
|
+
[
|
|
45
|
+
"rest",
|
|
46
|
+
"\\bcurl\\b[^\n]*https?://|\\bREST\\b|\\bopenapi\\b|/v1/"
|
|
47
|
+
],
|
|
48
|
+
[
|
|
49
|
+
"docker",
|
|
50
|
+
"(?i)\\bdocker\\b"
|
|
51
|
+
],
|
|
52
|
+
[
|
|
53
|
+
"plugin",
|
|
54
|
+
"(?i)/plugin install|\\.?claude-plugin|claude code plugin"
|
|
55
|
+
]
|
|
56
|
+
],
|
|
57
|
+
"min_stars": 10,
|
|
58
|
+
"days": 90
|
|
59
|
+
}
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "github-scrapers",
|
|
3
|
+
"description": "Tools that scrape, crawl or mine GitHub itself",
|
|
4
|
+
"queries": [
|
|
5
|
+
"topic:github-scraper",
|
|
6
|
+
"topic:github-crawler",
|
|
7
|
+
"topic:github-scraping",
|
|
8
|
+
"topic:github-mining",
|
|
9
|
+
"\"github scraper\" in:name,description",
|
|
10
|
+
"\"github crawler\" in:name,description",
|
|
11
|
+
"\"scrape github\" in:name,description",
|
|
12
|
+
"\"github scraping\" in:name,description",
|
|
13
|
+
"\"scrapes github\" in:description",
|
|
14
|
+
"\"crawl github\" in:name,description",
|
|
15
|
+
"\"github repository scraper\" in:description",
|
|
16
|
+
"\"github repo scraper\" in:description"
|
|
17
|
+
],
|
|
18
|
+
"exclude": "(?i)awesome|curated list|onlyfans",
|
|
19
|
+
"require": [
|
|
20
|
+
{
|
|
21
|
+
"pattern": "(?i)scrap|crawl|mining|miner|harvest"
|
|
22
|
+
},
|
|
23
|
+
{
|
|
24
|
+
"pattern": "(?i)github"
|
|
25
|
+
}
|
|
26
|
+
],
|
|
27
|
+
"interfaces": [
|
|
28
|
+
[
|
|
29
|
+
"cli",
|
|
30
|
+
"(?i)\\busage:|\\$ \\w+ --|\\bcommand line\\b|\\bcli\\b"
|
|
31
|
+
],
|
|
32
|
+
[
|
|
33
|
+
"pip",
|
|
34
|
+
"(?i)\\bpip3? install\\b|\\buv add\\b|\\bpypi\\b"
|
|
35
|
+
],
|
|
36
|
+
[
|
|
37
|
+
"npm",
|
|
38
|
+
"(?i)\\bnpm (i|install)\\b|\\bnpx\\b"
|
|
39
|
+
],
|
|
40
|
+
[
|
|
41
|
+
"action",
|
|
42
|
+
"(?i)uses: [\\w.-]+/[\\w.-]+@|github action"
|
|
43
|
+
],
|
|
44
|
+
[
|
|
45
|
+
"docker",
|
|
46
|
+
"(?i)\\bdocker\\b"
|
|
47
|
+
],
|
|
48
|
+
[
|
|
49
|
+
"token",
|
|
50
|
+
"(?i)GITHUB_TOKEN|personal access token"
|
|
51
|
+
]
|
|
52
|
+
],
|
|
53
|
+
"min_stars": 5,
|
|
54
|
+
"days": 365
|
|
55
|
+
}
|
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: topicscout
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional.
|
|
5
|
+
Author: Adecubed
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/adecubed/topicscout
|
|
8
|
+
Keywords: github,scraper,search,discovery,topics
|
|
9
|
+
Requires-Python: >=3.11
|
|
10
|
+
Description-Content-Type: text/markdown
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Dynamic: license-file
|
|
13
|
+
|
|
14
|
+
# topicscout
|
|
15
|
+
|
|
16
|
+
[](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
|
|
17
|
+
[](https://pypi.org/project/topicscout/)
|
|
18
|
+
|
|
19
|
+
Find the GitHub repositories on a topic, and next time only the new ones.
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
pip install topicscout # or run it without installing: uvx topicscout run github-scrapers
|
|
23
|
+
topicscout run github-scrapers
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
```
|
|
27
|
+
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
`new.md` is a table of what appeared since the last run, most stars first: stars,
|
|
31
|
+
last push, language, license, a guess at the interface from the README (MCP, pip,
|
|
32
|
+
npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
|
|
33
|
+
so a weekly run is a short list, not the same 300 repos again.
|
|
34
|
+
|
|
35
|
+
No dependencies, Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
|
|
36
|
+
are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.
|
|
37
|
+
|
|
38
|
+
## Why
|
|
39
|
+
|
|
40
|
+
GitHub search is good at "repos with this topic" and bad at "what's new since I last
|
|
41
|
+
looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
|
|
42
|
+
only tagged themselves". topicscout runs a fixed set of searches, applies your
|
|
43
|
+
filters, remembers what it has seen and what you already have.
|
|
44
|
+
|
|
45
|
+
It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
|
|
46
|
+
benchmark for AI memory systems, which uses it to find new memories to measure.
|
|
47
|
+
|
|
48
|
+
## Your own topic
|
|
49
|
+
|
|
50
|
+
Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
|
|
51
|
+
change the queries and the words a hit must contain, and run it:
|
|
52
|
+
|
|
53
|
+
```bash
|
|
54
|
+
topicscout run my-topic.json
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Run it again next week: `new.md` lists only what appeared in between. Put the repos
|
|
58
|
+
you already know or rejected in `scout-my-topic/known.txt`.
|
|
59
|
+
|
|
60
|
+
## From a script or an agent
|
|
61
|
+
|
|
62
|
+
`--json` prints the result on stdout, progress stays on stderr:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
topicscout run github-scrapers --json -q
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
```json
|
|
69
|
+
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
It is a command-line tool and a Python library (`from topicscout.core import Profile, run`),
|
|
73
|
+
not an MCP server; an agent calls it like any other command.
|
|
74
|
+
|
|
75
|
+
## Commands
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
|
|
79
|
+
[--known-from GLOB ...] [--json] [-q]
|
|
80
|
+
topicscout profiles # built-in profiles
|
|
81
|
+
topicscout doctor # token and remaining GitHub quota
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
- `PROFILE` is a built-in name or a path to your own JSON profile.
|
|
85
|
+
- `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
|
|
86
|
+
`new.md`, and your `known.txt`.
|
|
87
|
+
- `--json` prints the summary and the new repos as JSON on stdout; progress and
|
|
88
|
+
errors always go to stderr, so scripts and agents can read stdout as is.
|
|
89
|
+
|
|
90
|
+
## What counts as known
|
|
91
|
+
|
|
92
|
+
Repos you already have are marked `known` and never listed as new:
|
|
93
|
+
|
|
94
|
+
- `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
|
|
95
|
+
rejected repos here too.
|
|
96
|
+
- `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
|
|
97
|
+
files. adebench points it at its adapters, so a memory with an adapter drops out.
|
|
98
|
+
|
|
99
|
+
## Profiles
|
|
100
|
+
|
|
101
|
+
Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
|
|
102
|
+
(tools that scrape, crawl or mine GitHub). A profile is a JSON file:
|
|
103
|
+
|
|
104
|
+
```json
|
|
105
|
+
{
|
|
106
|
+
"name": "ai-memory",
|
|
107
|
+
"description": "Memory systems for AI agents and LLMs",
|
|
108
|
+
"queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
|
|
109
|
+
"exclude": "(?i)awesome|curated list|\\bpapers\\b",
|
|
110
|
+
"require": [
|
|
111
|
+
{"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
|
|
112
|
+
{"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
|
|
113
|
+
],
|
|
114
|
+
"interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
|
|
115
|
+
"min_stars": 10,
|
|
116
|
+
"days": 90
|
|
117
|
+
}
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
- `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
|
|
121
|
+
queries. Each one always gets `fork:false archived:false stars:>=min_stars
|
|
122
|
+
pushed:>=today-days`; up to 300 results per query.
|
|
123
|
+
- `exclude`: a regex on name, description and topics; a match drops the repo.
|
|
124
|
+
- `require`: every pattern must match. With `own_words` it must be in the name or
|
|
125
|
+
description, not only in a topic: a database that tagged itself `agent-memory`
|
|
126
|
+
is not a memory. Hyphens and underscores count as spaces.
|
|
127
|
+
- `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
|
|
128
|
+
Leave it out and no README is fetched.
|
|
129
|
+
- `require_language` (default true) drops repos with no language, which are
|
|
130
|
+
usually lists and docs.
|
|
131
|
+
|
|
132
|
+
Patterns are Python regexes; add `(?i)` for case-insensitive.
|
|
133
|
+
|
|
134
|
+
## Rate limits
|
|
135
|
+
|
|
136
|
+
On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
|
|
137
|
+
stops, saves what it found and says so in `new.md` and on stderr. The next run
|
|
138
|
+
picks up the READMEs it did not get to.
|
|
139
|
+
|
|
140
|
+
## How it compares
|
|
141
|
+
|
|
142
|
+
- [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
|
|
143
|
+
(issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
|
|
144
|
+
topic.
|
|
145
|
+
- [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
|
|
146
|
+
lists the top repos for a keyword, once; topicscout adds filters, profiles and
|
|
147
|
+
memory across runs.
|
|
148
|
+
|
|
149
|
+
## License
|
|
150
|
+
|
|
151
|
+
MIT
|
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
LICENSE
|
|
2
|
+
README.md
|
|
3
|
+
pyproject.toml
|
|
4
|
+
tests/test_topicscout.py
|
|
5
|
+
topicscout/__init__.py
|
|
6
|
+
topicscout/__main__.py
|
|
7
|
+
topicscout/cli.py
|
|
8
|
+
topicscout/core.py
|
|
9
|
+
topicscout.egg-info/PKG-INFO
|
|
10
|
+
topicscout.egg-info/SOURCES.txt
|
|
11
|
+
topicscout.egg-info/dependency_links.txt
|
|
12
|
+
topicscout.egg-info/entry_points.txt
|
|
13
|
+
topicscout.egg-info/top_level.txt
|
|
14
|
+
topicscout/profiles/ai-memory.json
|
|
15
|
+
topicscout/profiles/github-scrapers.json
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
topicscout
|