paper-scoring-digest 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- paper_scoring_digest-0.1.0/.gitignore +21 -0
- paper_scoring_digest-0.1.0/LICENSE +22 -0
- paper_scoring_digest-0.1.0/PKG-INFO +335 -0
- paper_scoring_digest-0.1.0/README.md +278 -0
- paper_scoring_digest-0.1.0/pyproject.toml +97 -0
- paper_scoring_digest-0.1.0/scripts/__init__.py +1 -0
- paper_scoring_digest-0.1.0/scripts/check_distribution.py +111 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/__init__.py +13 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/__main__.py +5 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/cli.py +307 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/discovery.py +137 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/models.py +66 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/presentation.py +159 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/py.typed +1 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/slack.py +89 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/storage.py +190 -0
- paper_scoring_digest-0.1.0/src/paper_scoring_digest/workflow.py +727 -0
- paper_scoring_digest-0.1.0/tests/conftest.py +32 -0
- paper_scoring_digest-0.1.0/tests/test_cli.py +285 -0
- paper_scoring_digest-0.1.0/tests/test_discovery.py +137 -0
- paper_scoring_digest-0.1.0/tests/test_distribution.py +78 -0
- paper_scoring_digest-0.1.0/tests/test_models.py +89 -0
- paper_scoring_digest-0.1.0/tests/test_presentation.py +149 -0
- paper_scoring_digest-0.1.0/tests/test_slack.py +262 -0
- paper_scoring_digest-0.1.0/tests/test_storage.py +325 -0
- paper_scoring_digest-0.1.0/tests/test_workflow.py +591 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
.env
|
|
2
|
+
.env.*
|
|
3
|
+
!.env.example
|
|
4
|
+
.venv/
|
|
5
|
+
__pycache__/
|
|
6
|
+
*.py[cod]
|
|
7
|
+
*.egg-info/
|
|
8
|
+
.coverage
|
|
9
|
+
coverage.xml
|
|
10
|
+
htmlcov/
|
|
11
|
+
.pytest_cache/
|
|
12
|
+
.mypy_cache/
|
|
13
|
+
.ruff_cache/
|
|
14
|
+
build/
|
|
15
|
+
dist/
|
|
16
|
+
reports/*.json
|
|
17
|
+
reports/*.jsonl
|
|
18
|
+
reports/*.sarif
|
|
19
|
+
!reports/*.xml
|
|
20
|
+
data/
|
|
21
|
+
*.log
|
|
@@ -0,0 +1,22 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Paper Scoring contributors
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
22
|
+
|
|
@@ -0,0 +1,335 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: paper-scoring-digest
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Daily ranked Paper Scoring digest for OpenClaw and Slack
|
|
5
|
+
Project-URL: Repository, https://github.com/maikeruSan/paper-scoring-digest
|
|
6
|
+
Author: Paper Scoring contributors
|
|
7
|
+
License: MIT License
|
|
8
|
+
|
|
9
|
+
Copyright (c) 2026 Paper Scoring contributors
|
|
10
|
+
|
|
11
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
12
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
13
|
+
in the Software without restriction, including without limitation the rights
|
|
14
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
15
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
16
|
+
furnished to do so, subject to the following conditions:
|
|
17
|
+
|
|
18
|
+
The above copyright notice and this permission notice shall be included in all
|
|
19
|
+
copies or substantial portions of the Software.
|
|
20
|
+
|
|
21
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
22
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
23
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
24
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
25
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
26
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
27
|
+
SOFTWARE.
|
|
28
|
+
|
|
29
|
+
License-File: LICENSE
|
|
30
|
+
Classifier: Development Status :: 3 - Alpha
|
|
31
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
32
|
+
Classifier: Programming Language :: Python :: 3
|
|
33
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
34
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
35
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
36
|
+
Requires-Python: >=3.11
|
|
37
|
+
Requires-Dist: paper-scoring-connectors[all]<0.2.0,>=0.1.0
|
|
38
|
+
Requires-Dist: paper-scoring-core<0.2.0,>=0.1.0
|
|
39
|
+
Requires-Dist: paper-scoring-pipeline<0.2.0,>=0.1.1
|
|
40
|
+
Requires-Dist: paper-scoring-reporting<0.2.0,>=0.1.0
|
|
41
|
+
Requires-Dist: pydantic<3,>=2.11
|
|
42
|
+
Requires-Dist: python-dotenv<2,>=1.1
|
|
43
|
+
Requires-Dist: tzdata>=2025.2; sys_platform == 'win32'
|
|
44
|
+
Provides-Extra: security
|
|
45
|
+
Requires-Dist: bandit[toml]<2,>=1.8; extra == 'security'
|
|
46
|
+
Requires-Dist: detect-secrets<2,>=1.5; extra == 'security'
|
|
47
|
+
Requires-Dist: pip-audit<3,>=2.9; extra == 'security'
|
|
48
|
+
Provides-Extra: test
|
|
49
|
+
Requires-Dist: build<2,>=1.3; extra == 'test'
|
|
50
|
+
Requires-Dist: coverage[toml]<8,>=7.10; extra == 'test'
|
|
51
|
+
Requires-Dist: mypy<2,>=1.18; extra == 'test'
|
|
52
|
+
Requires-Dist: packaging<27,>=25; extra == 'test'
|
|
53
|
+
Requires-Dist: pytest<10,>=9.0.3; extra == 'test'
|
|
54
|
+
Requires-Dist: ruff<1,>=0.12; extra == 'test'
|
|
55
|
+
Requires-Dist: twine<7,>=6.2; extra == 'test'
|
|
56
|
+
Description-Content-Type: text/markdown
|
|
57
|
+
|
|
58
|
+
# Paper Scoring Digest
|
|
59
|
+
|
|
60
|
+
`paper-scoring-digest` is the scheduling and delivery adapter for Paper
|
|
61
|
+
Scoring. It keeps arXiv discovery, local retention, Slack presentation, and
|
|
62
|
+
discussion state outside the provider-neutral core packages.
|
|
63
|
+
|
|
64
|
+
The daily workflow:
|
|
65
|
+
|
|
66
|
+
1. queries `cs.AI`, `cs.DB`, and `cs.LG` as one arXiv candidate pool;
|
|
67
|
+
2. removes cross-list duplicates before scoring;
|
|
68
|
+
3. ranks the merged pool and publishes one combined Top 20 (not 20 per
|
|
69
|
+
category);
|
|
70
|
+
4. downloads and retains the selected PDFs;
|
|
71
|
+
5. renders an email-like Slack card through OpenClaw; and
|
|
72
|
+
6. deletes PDFs that have reached seven days of age.
|
|
73
|
+
|
|
74
|
+
Metadata, scores, and arXiv links remain after the PDFs expire, so an Agent can
|
|
75
|
+
still explain an earlier ranking. A discussion that requires an expired PDF
|
|
76
|
+
must fetch it again from arXiv explicitly.
|
|
77
|
+
|
|
78
|
+
## Runtime contract
|
|
79
|
+
|
|
80
|
+
- Python 3.11, 3.12, or 3.13 on Ubuntu and Windows.
|
|
81
|
+
- Docker Engine or Docker Desktop with Docker Compose v2.
|
|
82
|
+
- Paper Scoring package-to-package dependencies come only from PyPI. Git,
|
|
83
|
+
local-path, and workspace dependency overrides are not used in production.
|
|
84
|
+
- Secrets are injected at process start. They are not copied into wheels,
|
|
85
|
+
container images, reports, GitHub Actions, or OpenClaw cron arguments.
|
|
86
|
+
- Scheduling and retention use the configured IANA timezone. Automatic arXiv
|
|
87
|
+
discovery selects the latest completed UTC submission date. The production
|
|
88
|
+
schedule below uses `Asia/Taipei`.
|
|
89
|
+
- Slack delivery uses a stable channel ID and a configured OpenClaw Slack
|
|
90
|
+
account. Slack tokens remain in OpenClaw's runtime secret store.
|
|
91
|
+
- On native Windows, collection, scoring, state inspection, and retention are
|
|
92
|
+
supported, but Slack delivery rejects npm `.cmd`/`.bat` shims because Windows
|
|
93
|
+
may parse them through `cmd.exe` despite `shell=False`. Run delivery from a
|
|
94
|
+
trusted Linux/WSL/Docker OpenClaw host or a native OpenClaw `.exe` launcher.
|
|
95
|
+
|
|
96
|
+
## Install from PyPI
|
|
97
|
+
|
|
98
|
+
Ubuntu:
|
|
99
|
+
|
|
100
|
+
```shell
|
|
101
|
+
python3 -m venv .venv
|
|
102
|
+
. .venv/bin/activate
|
|
103
|
+
python -m pip install --index-url https://pypi.org/simple "paper-scoring-digest==0.1.0"
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
Windows PowerShell:
|
|
107
|
+
|
|
108
|
+
```powershell
|
|
109
|
+
py -3.13 -m venv .venv
|
|
110
|
+
.\.venv\Scripts\Activate.ps1
|
|
111
|
+
python -m pip install --index-url https://pypi.org/simple "paper-scoring-digest==0.1.0"
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Copy `.env.example` to a location outside the repository, fill only the
|
|
115
|
+
provider values you need, and restrict access to that file. On Ubuntu:
|
|
116
|
+
|
|
117
|
+
```shell
|
|
118
|
+
install -d -m 700 "${HOME}/.config/paper-scoring"
|
|
119
|
+
install -m 600 .env.example "${HOME}/.config/paper-scoring/digest.env"
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
On Windows, store the file under a user-only directory and use the file's
|
|
123
|
+
Security properties or `icacls` to remove access for other users. Do not place
|
|
124
|
+
the file in the repository. Put the digest state directory under the same
|
|
125
|
+
user-only ACL: POSIX `0700`/`0600` mode changes do not configure Windows DACLs.
|
|
126
|
+
|
|
127
|
+
For an OpenAI-backed run, set `PAPER_SCORING_PROVIDER`,
|
|
128
|
+
`PAPER_SCORING_MODEL`, and `OPENAI_API_KEY`. For a local Ollama-backed run,
|
|
129
|
+
set the provider to `ollama`, select an installed model, and optionally set
|
|
130
|
+
`OLLAMA_BASE_URL`; no cloud API key is needed.
|
|
131
|
+
|
|
132
|
+
## Run once
|
|
133
|
+
|
|
134
|
+
The following is an illustrative production-shaped invocation. Replace every
|
|
135
|
+
angle-bracket placeholder; never paste a credential into the command line.
|
|
136
|
+
|
|
137
|
+
```shell
|
|
138
|
+
paper-scoring-digest run \
|
|
139
|
+
--env-file <ABSOLUTE_RUNTIME_ENV_FILE> \
|
|
140
|
+
--state-dir <ABSOLUTE_PRIVATE_STATE_DIRECTORY> \
|
|
141
|
+
--category cs.AI \
|
|
142
|
+
--category cs.DB \
|
|
143
|
+
--category cs.LG \
|
|
144
|
+
--top-n 20 \
|
|
145
|
+
--retention-days 7 \
|
|
146
|
+
--timezone Asia/Taipei \
|
|
147
|
+
--slack-channel <SLACK_CHANNEL_ID> \
|
|
148
|
+
--slack-account <OPENCLAW_SLACK_ACCOUNT>
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
By default the service pages through the complete daily category union before
|
|
152
|
+
ranking, so "Top 20" covers the full discovered day. The optional
|
|
153
|
+
`--candidate-limit 60` changes the meaning to "Top 20 from the latest 60
|
|
154
|
+
de-duplicated candidates" and bounds model cost; the Slack context labels that
|
|
155
|
+
scope explicitly. Re-running is idempotent: a pending delivery is handled
|
|
156
|
+
before a new daily run is created.
|
|
157
|
+
|
|
158
|
+
Useful read-only commands:
|
|
159
|
+
|
|
160
|
+
```shell
|
|
161
|
+
paper-scoring-digest show --run-id <YYYY-MM-DD> --rank 3
|
|
162
|
+
paper-scoring-digest show --run-id latest --paper-id <ARXIV_ID>
|
|
163
|
+
paper-scoring-digest presentation --run-id latest
|
|
164
|
+
paper-scoring-digest prune --state-dir <ABSOLUTE_PRIVATE_STATE_DIRECTORY> --retention-days 7 --timezone Asia/Taipei
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
Each successful run writes a private manifest plus text, HTML, and PDF reports
|
|
168
|
+
under the state directory. Only PDFs are removed by retention. Each manifest
|
|
169
|
+
stores `pdf_expires_at`, calculated from the instant when the digest was
|
|
170
|
+
created; an old arXiv submission downloaded during catch-up therefore still
|
|
171
|
+
receives the full 168-hour retention period.
|
|
172
|
+
|
|
173
|
+
## OpenClaw and Slack
|
|
174
|
+
|
|
175
|
+
Install the canonical cross-Agent Skill from PyPI, then install it into the
|
|
176
|
+
dedicated OpenClaw workspace:
|
|
177
|
+
|
|
178
|
+
```shell
|
|
179
|
+
python -m pip install --index-url https://pypi.org/simple "paper-scoring-skills[openai]>=0.2,<0.3"
|
|
180
|
+
paper-scoring-skills install --client openclaw --scope project
|
|
181
|
+
openclaw agents add paper-research \
|
|
182
|
+
--workspace <ABSOLUTE_OPENCLAW_AGENT_WORKSPACE> \
|
|
183
|
+
--model <PROVIDER/MODEL> \
|
|
184
|
+
--non-interactive
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
Configure the target Slack channel with `replyToMode: "all"` and add an exact
|
|
188
|
+
peer binding to `paper-research`. The fragment below is illustrative:
|
|
189
|
+
|
|
190
|
+
```json
|
|
191
|
+
{
|
|
192
|
+
"channels": {
|
|
193
|
+
"slack": {
|
|
194
|
+
"channels": {
|
|
195
|
+
"<SLACK_CHANNEL_ID>": { "replyToMode": "all" }
|
|
196
|
+
}
|
|
197
|
+
}
|
|
198
|
+
},
|
|
199
|
+
"bindings": [
|
|
200
|
+
{
|
|
201
|
+
"agentId": "paper-research",
|
|
202
|
+
"match": {
|
|
203
|
+
"channel": "slack",
|
|
204
|
+
"accountId": "<OPENCLAW_SLACK_ACCOUNT>",
|
|
205
|
+
"peer": { "kind": "channel", "id": "<SLACK_CHANNEL_ID>" }
|
|
206
|
+
}
|
|
207
|
+
}
|
|
208
|
+
]
|
|
209
|
+
}
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
`bindings` is an array and a config patch replaces arrays. Export and back up
|
|
213
|
+
the current OpenClaw config, merge this item with every existing binding, run
|
|
214
|
+
`openclaw config patch --dry-run`, and only then apply it. Do not paste the
|
|
215
|
+
fragment directly over a live bindings array.
|
|
216
|
+
|
|
217
|
+
Create the job disabled, test it, and enable it after the card and fallback
|
|
218
|
+
delivery have both been verified:
|
|
219
|
+
|
|
220
|
+
```shell
|
|
221
|
+
openclaw cron add \
|
|
222
|
+
--name paper-scoring-daily-digest \
|
|
223
|
+
--declaration-key paper-scoring-daily-digest-v1 \
|
|
224
|
+
--cron "15 19 * * *" \
|
|
225
|
+
--tz Asia/Taipei \
|
|
226
|
+
--exact \
|
|
227
|
+
--command-argv '["<VENV_BIN>/paper-scoring-digest","run","--env-file","<ABSOLUTE_RUNTIME_ENV_FILE>","--state-dir","<ABSOLUTE_PRIVATE_STATE_DIRECTORY>","--category","cs.AI","--category","cs.DB","--category","cs.LG","--top-n","20","--retention-days","7","--timezone","Asia/Taipei","--slack-channel","<SLACK_CHANNEL_ID>","--slack-account","<OPENCLAW_SLACK_ACCOUNT>","--cron-mode"]' \
|
|
228
|
+
--command-cwd <ABSOLUTE_OPENCLAW_AGENT_WORKSPACE> \
|
|
229
|
+
--timeout-seconds 14400 \
|
|
230
|
+
--announce \
|
|
231
|
+
--channel slack \
|
|
232
|
+
--account <OPENCLAW_SLACK_ACCOUNT> \
|
|
233
|
+
--to channel:<SLACK_CHANNEL_ID> \
|
|
234
|
+
--disabled
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
The digest sends a portable OpenClaw presentation that is rendered as Slack
|
|
238
|
+
Block Kit: a title, run context, one section per ranked paper, dividers, and a
|
|
239
|
+
thread instruction. The command's normal success output is suppressed in cron
|
|
240
|
+
mode; `--announce` is only a fallback if the command fails before direct card
|
|
241
|
+
delivery.
|
|
242
|
+
|
|
243
|
+
After the card arrives, reply in its thread, for example:
|
|
244
|
+
|
|
245
|
+
```text
|
|
246
|
+
3, 7:比較這兩篇的研究問題、方法與實驗結果
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
The exact channel binding routes the reply to `paper-research`. The Skill maps
|
|
250
|
+
the displayed rank to that card's stored run, reads a retained PDF when
|
|
251
|
+
available, and answers in the same thread. It does not silently re-score a
|
|
252
|
+
paper.
|
|
253
|
+
|
|
254
|
+
## Docker
|
|
255
|
+
|
|
256
|
+
The image builds the current package locally while resolving every Paper
|
|
257
|
+
Scoring dependency from the locked PyPI graph. The Dockerfile copies only the
|
|
258
|
+
package inputs; it never copies `.env` or the state directory.
|
|
259
|
+
|
|
260
|
+
Set the runtime env-file path in the shell that launches Compose:
|
|
261
|
+
|
|
262
|
+
```shell
|
|
263
|
+
export PAPER_SCORING_ENV_FILE="${HOME}/.config/paper-scoring/digest.env"
|
|
264
|
+
docker compose build
|
|
265
|
+
docker compose run --rm paper-scoring-digest --help
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
PowerShell:
|
|
269
|
+
|
|
270
|
+
```powershell
|
|
271
|
+
$env:PAPER_SCORING_ENV_FILE = "$HOME\.config\paper-scoring\digest.env"
|
|
272
|
+
docker compose build
|
|
273
|
+
docker compose run --rm paper-scoring-digest --help
|
|
274
|
+
```
|
|
275
|
+
|
|
276
|
+
The Compose service uses a named volume for `/data`, which avoids host
|
|
277
|
+
ownership differences between Ubuntu and Docker Desktop. Provider credentials
|
|
278
|
+
are loaded only through `env_file` when the container starts.
|
|
279
|
+
|
|
280
|
+
To collect, score, and persist a digest in the container without attempting to
|
|
281
|
+
access the host's OpenClaw CLI, use `--no-deliver`:
|
|
282
|
+
|
|
283
|
+
```shell
|
|
284
|
+
docker compose run --rm paper-scoring-digest run \
|
|
285
|
+
--no-deliver \
|
|
286
|
+
--state-dir /data \
|
|
287
|
+
--category cs.AI --category cs.DB --category cs.LG \
|
|
288
|
+
--top-n 20 --retention-days 7
|
|
289
|
+
```
|
|
290
|
+
|
|
291
|
+
The manifest remains `pending`; a trusted OpenClaw host with access to the
|
|
292
|
+
same state directory can later run the normal delivery command without
|
|
293
|
+
rescoring.
|
|
294
|
+
|
|
295
|
+
The Python image intentionally does not bundle OpenClaw or a Slack token. Use
|
|
296
|
+
it for collection, ranking, report/state inspection, and retention inside a
|
|
297
|
+
container; run Slack delivery from the trusted OpenClaw host (or extend the
|
|
298
|
+
official OpenClaw image with this PyPI package). This keeps the Gateway's
|
|
299
|
+
credentials and config out of the digest image.
|
|
300
|
+
|
|
301
|
+
## Verification and release evidence
|
|
302
|
+
|
|
303
|
+
Pull requests and `main` run the following gates:
|
|
304
|
+
|
|
305
|
+
- Ubuntu and Windows tests on Python 3.11-3.13, with branch coverage and JUnit
|
|
306
|
+
reports;
|
|
307
|
+
- Ruff, formatting, strict mypy, wheel/sdist build, and Twine validation;
|
|
308
|
+
- a non-root, read-only Docker smoke test;
|
|
309
|
+
- Bandit, `pip-audit`, and `detect-secrets`; and
|
|
310
|
+
- a redacted Gitleaks scan of the complete Git history plus a noreply-author
|
|
311
|
+
policy check.
|
|
312
|
+
|
|
313
|
+
The workflows upload test, package, and security reports as GitHub Actions
|
|
314
|
+
artifacts even when a gate fails. Reports must be reviewed before merging,
|
|
315
|
+
tagging, or publishing to PyPI. Artifacts are evidence, not a place to store
|
|
316
|
+
runtime credentials.
|
|
317
|
+
|
|
318
|
+
For a local equivalent using only the locked PyPI dependency graph:
|
|
319
|
+
|
|
320
|
+
```shell
|
|
321
|
+
uv sync --locked --no-sources --extra test --extra security
|
|
322
|
+
uv run --frozen ruff check .
|
|
323
|
+
uv run --frozen ruff format --check .
|
|
324
|
+
uv run --frozen mypy
|
|
325
|
+
uv run --frozen coverage run --branch -m pytest
|
|
326
|
+
uv run --frozen coverage report --show-missing
|
|
327
|
+
uv run --frozen bandit -c pyproject.toml -r src
|
|
328
|
+
uv run --frozen pip-audit
|
|
329
|
+
uv build --no-sources --default-index https://pypi.org/simple
|
|
330
|
+
uv run --frozen twine check dist/*
|
|
331
|
+
```
|
|
332
|
+
|
|
333
|
+
Before a release, also scan untracked files and the complete Git history. Never
|
|
334
|
+
publish when any report contains an unexplained secret candidate, personal
|
|
335
|
+
path, private email address, channel ID, token, or credential.
|
|
@@ -0,0 +1,278 @@
|
|
|
1
|
+
# Paper Scoring Digest
|
|
2
|
+
|
|
3
|
+
`paper-scoring-digest` is the scheduling and delivery adapter for Paper
|
|
4
|
+
Scoring. It keeps arXiv discovery, local retention, Slack presentation, and
|
|
5
|
+
discussion state outside the provider-neutral core packages.
|
|
6
|
+
|
|
7
|
+
The daily workflow:
|
|
8
|
+
|
|
9
|
+
1. queries `cs.AI`, `cs.DB`, and `cs.LG` as one arXiv candidate pool;
|
|
10
|
+
2. removes cross-list duplicates before scoring;
|
|
11
|
+
3. ranks the merged pool and publishes one combined Top 20 (not 20 per
|
|
12
|
+
category);
|
|
13
|
+
4. downloads and retains the selected PDFs;
|
|
14
|
+
5. renders an email-like Slack card through OpenClaw; and
|
|
15
|
+
6. deletes PDFs that have reached seven days of age.
|
|
16
|
+
|
|
17
|
+
Metadata, scores, and arXiv links remain after the PDFs expire, so an Agent can
|
|
18
|
+
still explain an earlier ranking. A discussion that requires an expired PDF
|
|
19
|
+
must fetch it again from arXiv explicitly.
|
|
20
|
+
|
|
21
|
+
## Runtime contract
|
|
22
|
+
|
|
23
|
+
- Python 3.11, 3.12, or 3.13 on Ubuntu and Windows.
|
|
24
|
+
- Docker Engine or Docker Desktop with Docker Compose v2.
|
|
25
|
+
- Paper Scoring package-to-package dependencies come only from PyPI. Git,
|
|
26
|
+
local-path, and workspace dependency overrides are not used in production.
|
|
27
|
+
- Secrets are injected at process start. They are not copied into wheels,
|
|
28
|
+
container images, reports, GitHub Actions, or OpenClaw cron arguments.
|
|
29
|
+
- Scheduling and retention use the configured IANA timezone. Automatic arXiv
|
|
30
|
+
discovery selects the latest completed UTC submission date. The production
|
|
31
|
+
schedule below uses `Asia/Taipei`.
|
|
32
|
+
- Slack delivery uses a stable channel ID and a configured OpenClaw Slack
|
|
33
|
+
account. Slack tokens remain in OpenClaw's runtime secret store.
|
|
34
|
+
- On native Windows, collection, scoring, state inspection, and retention are
|
|
35
|
+
supported, but Slack delivery rejects npm `.cmd`/`.bat` shims because Windows
|
|
36
|
+
may parse them through `cmd.exe` despite `shell=False`. Run delivery from a
|
|
37
|
+
trusted Linux/WSL/Docker OpenClaw host or a native OpenClaw `.exe` launcher.
|
|
38
|
+
|
|
39
|
+
## Install from PyPI
|
|
40
|
+
|
|
41
|
+
Ubuntu:
|
|
42
|
+
|
|
43
|
+
```shell
|
|
44
|
+
python3 -m venv .venv
|
|
45
|
+
. .venv/bin/activate
|
|
46
|
+
python -m pip install --index-url https://pypi.org/simple "paper-scoring-digest==0.1.0"
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
Windows PowerShell:
|
|
50
|
+
|
|
51
|
+
```powershell
|
|
52
|
+
py -3.13 -m venv .venv
|
|
53
|
+
.\.venv\Scripts\Activate.ps1
|
|
54
|
+
python -m pip install --index-url https://pypi.org/simple "paper-scoring-digest==0.1.0"
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Copy `.env.example` to a location outside the repository, fill only the
|
|
58
|
+
provider values you need, and restrict access to that file. On Ubuntu:
|
|
59
|
+
|
|
60
|
+
```shell
|
|
61
|
+
install -d -m 700 "${HOME}/.config/paper-scoring"
|
|
62
|
+
install -m 600 .env.example "${HOME}/.config/paper-scoring/digest.env"
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
On Windows, store the file under a user-only directory and use the file's
|
|
66
|
+
Security properties or `icacls` to remove access for other users. Do not place
|
|
67
|
+
the file in the repository. Put the digest state directory under the same
|
|
68
|
+
user-only ACL: POSIX `0700`/`0600` mode changes do not configure Windows DACLs.
|
|
69
|
+
|
|
70
|
+
For an OpenAI-backed run, set `PAPER_SCORING_PROVIDER`,
|
|
71
|
+
`PAPER_SCORING_MODEL`, and `OPENAI_API_KEY`. For a local Ollama-backed run,
|
|
72
|
+
set the provider to `ollama`, select an installed model, and optionally set
|
|
73
|
+
`OLLAMA_BASE_URL`; no cloud API key is needed.
|
|
74
|
+
|
|
75
|
+
## Run once
|
|
76
|
+
|
|
77
|
+
The following is an illustrative production-shaped invocation. Replace every
|
|
78
|
+
angle-bracket placeholder; never paste a credential into the command line.
|
|
79
|
+
|
|
80
|
+
```shell
|
|
81
|
+
paper-scoring-digest run \
|
|
82
|
+
--env-file <ABSOLUTE_RUNTIME_ENV_FILE> \
|
|
83
|
+
--state-dir <ABSOLUTE_PRIVATE_STATE_DIRECTORY> \
|
|
84
|
+
--category cs.AI \
|
|
85
|
+
--category cs.DB \
|
|
86
|
+
--category cs.LG \
|
|
87
|
+
--top-n 20 \
|
|
88
|
+
--retention-days 7 \
|
|
89
|
+
--timezone Asia/Taipei \
|
|
90
|
+
--slack-channel <SLACK_CHANNEL_ID> \
|
|
91
|
+
--slack-account <OPENCLAW_SLACK_ACCOUNT>
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
By default the service pages through the complete daily category union before
|
|
95
|
+
ranking, so "Top 20" covers the full discovered day. The optional
|
|
96
|
+
`--candidate-limit 60` changes the meaning to "Top 20 from the latest 60
|
|
97
|
+
de-duplicated candidates" and bounds model cost; the Slack context labels that
|
|
98
|
+
scope explicitly. Re-running is idempotent: a pending delivery is handled
|
|
99
|
+
before a new daily run is created.
|
|
100
|
+
|
|
101
|
+
Useful read-only commands:
|
|
102
|
+
|
|
103
|
+
```shell
|
|
104
|
+
paper-scoring-digest show --run-id <YYYY-MM-DD> --rank 3
|
|
105
|
+
paper-scoring-digest show --run-id latest --paper-id <ARXIV_ID>
|
|
106
|
+
paper-scoring-digest presentation --run-id latest
|
|
107
|
+
paper-scoring-digest prune --state-dir <ABSOLUTE_PRIVATE_STATE_DIRECTORY> --retention-days 7 --timezone Asia/Taipei
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
Each successful run writes a private manifest plus text, HTML, and PDF reports
|
|
111
|
+
under the state directory. Only PDFs are removed by retention. Each manifest
|
|
112
|
+
stores `pdf_expires_at`, calculated from the instant when the digest was
|
|
113
|
+
created; an old arXiv submission downloaded during catch-up therefore still
|
|
114
|
+
receives the full 168-hour retention period.
|
|
115
|
+
|
|
116
|
+
## OpenClaw and Slack
|
|
117
|
+
|
|
118
|
+
Install the canonical cross-Agent Skill from PyPI, then install it into the
|
|
119
|
+
dedicated OpenClaw workspace:
|
|
120
|
+
|
|
121
|
+
```shell
|
|
122
|
+
python -m pip install --index-url https://pypi.org/simple "paper-scoring-skills[openai]>=0.2,<0.3"
|
|
123
|
+
paper-scoring-skills install --client openclaw --scope project
|
|
124
|
+
openclaw agents add paper-research \
|
|
125
|
+
--workspace <ABSOLUTE_OPENCLAW_AGENT_WORKSPACE> \
|
|
126
|
+
--model <PROVIDER/MODEL> \
|
|
127
|
+
--non-interactive
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
Configure the target Slack channel with `replyToMode: "all"` and add an exact
|
|
131
|
+
peer binding to `paper-research`. The fragment below is illustrative:
|
|
132
|
+
|
|
133
|
+
```json
|
|
134
|
+
{
|
|
135
|
+
"channels": {
|
|
136
|
+
"slack": {
|
|
137
|
+
"channels": {
|
|
138
|
+
"<SLACK_CHANNEL_ID>": { "replyToMode": "all" }
|
|
139
|
+
}
|
|
140
|
+
}
|
|
141
|
+
},
|
|
142
|
+
"bindings": [
|
|
143
|
+
{
|
|
144
|
+
"agentId": "paper-research",
|
|
145
|
+
"match": {
|
|
146
|
+
"channel": "slack",
|
|
147
|
+
"accountId": "<OPENCLAW_SLACK_ACCOUNT>",
|
|
148
|
+
"peer": { "kind": "channel", "id": "<SLACK_CHANNEL_ID>" }
|
|
149
|
+
}
|
|
150
|
+
}
|
|
151
|
+
]
|
|
152
|
+
}
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
`bindings` is an array and a config patch replaces arrays. Export and back up
|
|
156
|
+
the current OpenClaw config, merge this item with every existing binding, run
|
|
157
|
+
`openclaw config patch --dry-run`, and only then apply it. Do not paste the
|
|
158
|
+
fragment directly over a live bindings array.
|
|
159
|
+
|
|
160
|
+
Create the job disabled, test it, and enable it after the card and fallback
|
|
161
|
+
delivery have both been verified:
|
|
162
|
+
|
|
163
|
+
```shell
|
|
164
|
+
openclaw cron add \
|
|
165
|
+
--name paper-scoring-daily-digest \
|
|
166
|
+
--declaration-key paper-scoring-daily-digest-v1 \
|
|
167
|
+
--cron "15 19 * * *" \
|
|
168
|
+
--tz Asia/Taipei \
|
|
169
|
+
--exact \
|
|
170
|
+
--command-argv '["<VENV_BIN>/paper-scoring-digest","run","--env-file","<ABSOLUTE_RUNTIME_ENV_FILE>","--state-dir","<ABSOLUTE_PRIVATE_STATE_DIRECTORY>","--category","cs.AI","--category","cs.DB","--category","cs.LG","--top-n","20","--retention-days","7","--timezone","Asia/Taipei","--slack-channel","<SLACK_CHANNEL_ID>","--slack-account","<OPENCLAW_SLACK_ACCOUNT>","--cron-mode"]' \
|
|
171
|
+
--command-cwd <ABSOLUTE_OPENCLAW_AGENT_WORKSPACE> \
|
|
172
|
+
--timeout-seconds 14400 \
|
|
173
|
+
--announce \
|
|
174
|
+
--channel slack \
|
|
175
|
+
--account <OPENCLAW_SLACK_ACCOUNT> \
|
|
176
|
+
--to channel:<SLACK_CHANNEL_ID> \
|
|
177
|
+
--disabled
|
|
178
|
+
```
|
|
179
|
+
|
|
180
|
+
The digest sends a portable OpenClaw presentation that is rendered as Slack
|
|
181
|
+
Block Kit: a title, run context, one section per ranked paper, dividers, and a
|
|
182
|
+
thread instruction. The command's normal success output is suppressed in cron
|
|
183
|
+
mode; `--announce` is only a fallback if the command fails before direct card
|
|
184
|
+
delivery.
|
|
185
|
+
|
|
186
|
+
After the card arrives, reply in its thread, for example:
|
|
187
|
+
|
|
188
|
+
```text
|
|
189
|
+
3, 7:比較這兩篇的研究問題、方法與實驗結果
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
The exact channel binding routes the reply to `paper-research`. The Skill maps
|
|
193
|
+
the displayed rank to that card's stored run, reads a retained PDF when
|
|
194
|
+
available, and answers in the same thread. It does not silently re-score a
|
|
195
|
+
paper.
|
|
196
|
+
|
|
197
|
+
## Docker
|
|
198
|
+
|
|
199
|
+
The image builds the current package locally while resolving every Paper
|
|
200
|
+
Scoring dependency from the locked PyPI graph. The Dockerfile copies only the
|
|
201
|
+
package inputs; it never copies `.env` or the state directory.
|
|
202
|
+
|
|
203
|
+
Set the runtime env-file path in the shell that launches Compose:
|
|
204
|
+
|
|
205
|
+
```shell
|
|
206
|
+
export PAPER_SCORING_ENV_FILE="${HOME}/.config/paper-scoring/digest.env"
|
|
207
|
+
docker compose build
|
|
208
|
+
docker compose run --rm paper-scoring-digest --help
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
PowerShell:
|
|
212
|
+
|
|
213
|
+
```powershell
|
|
214
|
+
$env:PAPER_SCORING_ENV_FILE = "$HOME\.config\paper-scoring\digest.env"
|
|
215
|
+
docker compose build
|
|
216
|
+
docker compose run --rm paper-scoring-digest --help
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
The Compose service uses a named volume for `/data`, which avoids host
|
|
220
|
+
ownership differences between Ubuntu and Docker Desktop. Provider credentials
|
|
221
|
+
are loaded only through `env_file` when the container starts.
|
|
222
|
+
|
|
223
|
+
To collect, score, and persist a digest in the container without attempting to
|
|
224
|
+
access the host's OpenClaw CLI, use `--no-deliver`:
|
|
225
|
+
|
|
226
|
+
```shell
|
|
227
|
+
docker compose run --rm paper-scoring-digest run \
|
|
228
|
+
--no-deliver \
|
|
229
|
+
--state-dir /data \
|
|
230
|
+
--category cs.AI --category cs.DB --category cs.LG \
|
|
231
|
+
--top-n 20 --retention-days 7
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
The manifest remains `pending`; a trusted OpenClaw host with access to the
|
|
235
|
+
same state directory can later run the normal delivery command without
|
|
236
|
+
rescoring.
|
|
237
|
+
|
|
238
|
+
The Python image intentionally does not bundle OpenClaw or a Slack token. Use
|
|
239
|
+
it for collection, ranking, report/state inspection, and retention inside a
|
|
240
|
+
container; run Slack delivery from the trusted OpenClaw host (or extend the
|
|
241
|
+
official OpenClaw image with this PyPI package). This keeps the Gateway's
|
|
242
|
+
credentials and config out of the digest image.
|
|
243
|
+
|
|
244
|
+
## Verification and release evidence
|
|
245
|
+
|
|
246
|
+
Pull requests and `main` run the following gates:
|
|
247
|
+
|
|
248
|
+
- Ubuntu and Windows tests on Python 3.11-3.13, with branch coverage and JUnit
|
|
249
|
+
reports;
|
|
250
|
+
- Ruff, formatting, strict mypy, wheel/sdist build, and Twine validation;
|
|
251
|
+
- a non-root, read-only Docker smoke test;
|
|
252
|
+
- Bandit, `pip-audit`, and `detect-secrets`; and
|
|
253
|
+
- a redacted Gitleaks scan of the complete Git history plus a noreply-author
|
|
254
|
+
policy check.
|
|
255
|
+
|
|
256
|
+
The workflows upload test, package, and security reports as GitHub Actions
|
|
257
|
+
artifacts even when a gate fails. Reports must be reviewed before merging,
|
|
258
|
+
tagging, or publishing to PyPI. Artifacts are evidence, not a place to store
|
|
259
|
+
runtime credentials.
|
|
260
|
+
|
|
261
|
+
For a local equivalent using only the locked PyPI dependency graph:
|
|
262
|
+
|
|
263
|
+
```shell
|
|
264
|
+
uv sync --locked --no-sources --extra test --extra security
|
|
265
|
+
uv run --frozen ruff check .
|
|
266
|
+
uv run --frozen ruff format --check .
|
|
267
|
+
uv run --frozen mypy
|
|
268
|
+
uv run --frozen coverage run --branch -m pytest
|
|
269
|
+
uv run --frozen coverage report --show-missing
|
|
270
|
+
uv run --frozen bandit -c pyproject.toml -r src
|
|
271
|
+
uv run --frozen pip-audit
|
|
272
|
+
uv build --no-sources --default-index https://pypi.org/simple
|
|
273
|
+
uv run --frozen twine check dist/*
|
|
274
|
+
```
|
|
275
|
+
|
|
276
|
+
Before a release, also scan untracked files and the complete Git history. Never
|
|
277
|
+
publish when any report contains an unexplained secret candidate, personal
|
|
278
|
+
path, private email address, channel ID, token, or credential.
|