cronwatch-sdk 0.6.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- cronwatch_sdk-0.6.0/.gitignore +10 -0
- cronwatch_sdk-0.6.0/DESIGN.md +272 -0
- cronwatch_sdk-0.6.0/LICENSE +21 -0
- cronwatch_sdk-0.6.0/PKG-INFO +259 -0
- cronwatch_sdk-0.6.0/README.md +215 -0
- cronwatch_sdk-0.6.0/pyproject.toml +76 -0
- cronwatch_sdk-0.6.0/src/cronwatch/__init__.py +115 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_convert.py +98 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_cron.py +524 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_env.py +43 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_js.py +280 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_response.py +68 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_scheduler_check.py +180 -0
- cronwatch_sdk-0.6.0/src/cronwatch/_zone.py +96 -0
- cronwatch_sdk-0.6.0/src/cronwatch/aio.py +223 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/__init__.py +147 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/_http.py +175 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/_shared.py +218 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/bugsnag.py +84 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/datadog.py +83 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/discord.py +67 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/email.py +93 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/honeybadger.py +83 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/mailgun.py +51 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/newrelic.py +72 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/postmark.py +57 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/resend.py +53 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/rollbar.py +68 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/sendgrid.py +54 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/sentry.py +117 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/ses.py +94 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/sigv4.py +96 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/slack.py +70 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/twilio.py +219 -0
- cronwatch_sdk-0.6.0/src/cronwatch/alerts/webhook.py +51 -0
- cronwatch_sdk-0.6.0/src/cronwatch/apscheduler.py +520 -0
- cronwatch_sdk-0.6.0/src/cronwatch/celery.py +765 -0
- cronwatch_sdk-0.6.0/src/cronwatch/client.py +1591 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/__init__.py +176 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/apps.py +34 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/management/__init__.py +0 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/management/commands/__init__.py +0 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/management/commands/cronwatch_check.py +27 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/urls.py +16 -0
- cronwatch_sdk-0.6.0/src/cronwatch/django/views.py +55 -0
- cronwatch_sdk-0.6.0/src/cronwatch/duration.py +70 -0
- cronwatch_sdk-0.6.0/src/cronwatch/evaluate.py +388 -0
- cronwatch_sdk-0.6.0/src/cronwatch/format.py +120 -0
- cronwatch_sdk-0.6.0/src/cronwatch/handler.py +461 -0
- cronwatch_sdk-0.6.0/src/cronwatch/job.py +163 -0
- cronwatch_sdk-0.6.0/src/cronwatch/output.py +214 -0
- cronwatch_sdk-0.6.0/src/cronwatch/py.typed +0 -0
- cronwatch_sdk-0.6.0/src/cronwatch/run_handle.py +214 -0
- cronwatch_sdk-0.6.0/src/cronwatch/schedule.py +221 -0
- cronwatch_sdk-0.6.0/src/cronwatch/serialize.py +68 -0
- cronwatch_sdk-0.6.0/src/cronwatch/sources/__init__.py +8 -0
- cronwatch_sdk-0.6.0/src/cronwatch/sources/pgcron.py +560 -0
- cronwatch_sdk-0.6.0/src/cronwatch/stats.py +22 -0
- cronwatch_sdk-0.6.0/src/cronwatch/stores/__init__.py +39 -0
- cronwatch_sdk-0.6.0/src/cronwatch/stores/_sql.py +211 -0
- cronwatch_sdk-0.6.0/src/cronwatch/stores/memory.py +166 -0
- cronwatch_sdk-0.6.0/src/cronwatch/stores/postgres.py +180 -0
- cronwatch_sdk-0.6.0/src/cronwatch/stores/sqlite.py +201 -0
- cronwatch_sdk-0.6.0/src/cronwatch/triage/__init__.py +3 -0
- cronwatch_sdk-0.6.0/src/cronwatch/triage/anthropic.py +170 -0
- cronwatch_sdk-0.6.0/src/cronwatch/types.py +492 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/__init__.py +891 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/_escape.py +71 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/_html.py +538 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/_icons.py +12 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/_origin.py +176 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/_pwa.py +147 -0
- cronwatch_sdk-0.6.0/src/cronwatch/web/_timeline.py +536 -0
|
@@ -0,0 +1,272 @@
|
|
|
1
|
+
# The Python port
|
|
2
|
+
|
|
3
|
+
`packages/python` is `cronwatch-sdk` on PyPI (import name `cronwatch`; the name `cronwatch` on PyPI belongs to an abandoned project): the same library as `@cronwatch/sdk`, for Python apps. It is a port, not a new design, made the way the Ruby gem was (see `packages/ruby/DESIGN.md`). The TypeScript SDK is the source of truth for every behaviour, message and stored byte; when the two disagree, the Python side is wrong.
|
|
4
|
+
|
|
5
|
+
## Rules
|
|
6
|
+
|
|
7
|
+
- Python 3.11 or newer (`enum.StrEnum`, `zoneinfo`, `sqlite3` error names). The core has no runtime dependencies. Cron parsing and fire times are a port of croner 10 (`_cron.py`), zones come from the standard library's `zoneinfo`, and the SQLite store uses the standard library's `sqlite3`. On Windows, which has no zone database, `pip install "cronwatch-sdk[tzdata]"`.
|
|
8
|
+
- Everything else is optional and loaded only from its own module, each raising an `ImportError` that names the package to add when it is missing: `cronwatch.stores.postgres` (psycopg 3.2 or newer, the extra `cronwatch-sdk[postgres]`), `cronwatch.triage.anthropic` (anthropic, the extra `cronwatch-sdk[anthropic]`), `cronwatch.django` (Django 5.2 or newer, the extra `cronwatch-sdk[django]`), `cronwatch.celery` (Celery 5.5 or newer, the extra `cronwatch-sdk[celery]`; django-celery-beat is read when installed, never required) and `cronwatch.apscheduler` (APScheduler 3.10 or newer in the 3 line, the extra `cronwatch-sdk[apscheduler]`). `cronwatch.aio` and `cronwatch.handler` use the standard library only, and a handler imports Django, Werkzeug or Starlette only to answer a request of that framework. The dashboard (`cronwatch.web`) uses the standard library only. The alert channels use the standard library only and load with `cronwatch.alerts`; the pg_cron source (`cronwatch.sources.pgcron`) needs psycopg only when given a psycopg connection, pool or connection string.
|
|
9
|
+
- Times are `int` epoch milliseconds everywhere, as in the SDK and the gem, so `evaluate` ports line for line and stored rows are identical.
|
|
10
|
+
- Python names are snake_case (`failures_before_alert`, `max_duration`, `started_at`). Run statuses, conditions, alert types and health are `StrEnum`s whose values are the wire strings (`Condition.OVER_BUDGET == "over_budget"`), so they compare equal to plain strings, hash like them and print as them. Anything that leaves the process (store rows, JSON columns, alert JSON) uses the SDK's exact camelCase field names, key order and string values, so a Node, a Ruby and a Python process can share one database and `@cronwatch/mcp` works against any of them.
|
|
11
|
+
- JSON is written by `_js.dumps`, which is `JSON.stringify` byte for byte: numbers as JavaScript prints them (`2` not `2.0`, `1e-7`, `1e+21`), integer-like keys first in ascending order and the rest as inserted, only control characters, quotes, backslashes and lone surrogates escaped. `to_dict()` on every type is its JSON shape in the SDK's key order; `from_dict()` reads camelCase or snake_case keys.
|
|
12
|
+
- Alert titles and messages are the SDK's text, character for character. `JobDefinition` keeps its fields as given, in order (defaults, then options as given, then `name`, and a stored `expect` last), including fields a newer writer added, so the stored definition is the SDK's JSON.
|
|
13
|
+
- Lengths and cuts are in UTF-16 code units, as JavaScript counts them (`_js.length16`, `head16`, `tail16`). A cut through a surrogate pair leaves U+FFFD where JavaScript would keep a lone surrogate, which is the character that surrogate becomes once written out as UTF-8 (to a store, a hash, a network), so the stored bytes are the same.
|
|
14
|
+
- Secret redaction uses the SDK's patterns, written so Python's `re` matches exactly what JavaScript's matches: `re.ASCII` keeps `\b` and the classes to ASCII, JavaScript's `\s` is spelled out, case-insensitive words are spelled `[Ss][Ee]...` (ASCII-only folding, as JavaScript's `/i` has it here), and text with characters outside the Basic Multilingual Plane is matched as surrogate pairs, then put back together.
|
|
15
|
+
- Every condition opens once and closes with a recovery. No repeat alerts. A job whose schedule is removed while missed is open has missed closed by the next check with a recovery of its own (`reason: "unscheduled"`).
|
|
16
|
+
- A job's function may raise anything. An `Exception` is recorded as the failure and raised again; anything outside it (`KeyboardInterrupt`, `SystemExit`, `GeneratorExit`) is recorded as a failed run (`Interrupted: KeyboardInterrupt`) and raised again, so no run is left `running` to be reported stuck later. The store failing never stops a job: store errors go to `on_error`, and the job's own outcome is returned or raised.
|
|
17
|
+
- An error is written `Name: message` and up to five frames, innermost first, each ` at function (file:line)`, as a JavaScript stack reads.
|
|
18
|
+
- The environment is read in one place, `_env.py`: the first of `CRONWATCH_ENV`, `APP_ENV` and `ENVIRONMENT` that is set (the SDK reads `NODE_ENV`), else what a framework integration says (`cronwatch.django` says development when `DEBUG` is on and production when it is off). It decides whether the in-memory store warns that it forgets on restart, and whether the dashboard makes a development token when none is configured.
|
|
19
|
+
- A client survives a fork (gunicorn, Celery's prefork pool): it notes its pid, and in a child it starts with fresh locks, no check in flight and no interval thread, so `start()` and `check()` work there.
|
|
20
|
+
- No em or en dashes anywhere, as in the rest of the repo.
|
|
21
|
+
|
|
22
|
+
## Layout
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
packages/python/
|
|
26
|
+
pyproject.toml README.md LICENSE DESIGN.md
|
|
27
|
+
src/cronwatch/
|
|
28
|
+
__init__.py Cronwatch, configure() and client(), current(), the public names, __version__
|
|
29
|
+
py.typed
|
|
30
|
+
_js.py JavaScript's numbers, JSON.stringify, trim, \s, UTF-16 lengths, Date.UTC and toISOString
|
|
31
|
+
_env.py the environment
|
|
32
|
+
_zone.py zoneinfo zones (case-insensitive, as Intl), croner's wall-clock arithmetic (fromTZ)
|
|
33
|
+
_cron.py croner: CronPattern (the reading of an expression, its checks and messages) and CronDate (the walk)
|
|
34
|
+
types.py Run, JobState, StoredJob, JobDefinition, Alert, AlertDraft, JobSummary, CheckResult, the StrEnums
|
|
35
|
+
duration.py "15m", "1h30m", timedelta: parse and format
|
|
36
|
+
stats.py percentiles
|
|
37
|
+
output.py the output cap, error messages, secret redaction
|
|
38
|
+
schedule.py parse_schedule, next_fire, expectation, run_covers (schedule.ts)
|
|
39
|
+
evaluate.py the alert rules, pure functions (evaluate.ts)
|
|
40
|
+
format.py alert titles and messages
|
|
41
|
+
serialize.py stored definitions, expect rules
|
|
42
|
+
job.py JobContext (log, metric, signal), RunRecorder, AbortSignal, current()
|
|
43
|
+
run_handle.py RunHandle: a run started by job.start() or found by job.resume(), finished later
|
|
44
|
+
alerts/
|
|
45
|
+
__init__.py the channel protocol, ChannelContext, Console, Custom, and every channel's class
|
|
46
|
+
_http.py the POSTs, on urllib.request: no redirect followed, one ten second deadline (Response, UrllibHTTP)
|
|
47
|
+
_shared.py alerts/shared.ts: the POST with redacted errors, alert ids, run summaries, JavaScript's cuts and encodings
|
|
48
|
+
email.py subject, text and HTML for every email channel (alerts/email.ts)
|
|
49
|
+
sigv4.py AWS Signature Version 4 on hashlib and hmac (alerts/sigv4.ts)
|
|
50
|
+
slack.py discord.py webhook.py resend.py postmark.py sendgrid.py mailgun.py ses.py twilio.py
|
|
51
|
+
sentry.py honeybadger.py datadog.py rollbar.py bugsnag.py newrelic.py
|
|
52
|
+
sources/
|
|
53
|
+
pgcron.py PgCron, the pg_cron reader (sources/pgcron.ts), and its psycopg adapter
|
|
54
|
+
triage/
|
|
55
|
+
anthropic.py AnthropicTriage, Claude through the anthropic package (triage/anthropic.ts)
|
|
56
|
+
client.py Cronwatch and JobHandle: job, run (plain and async), check, start/stop, silence, forget, jobs, runs, record_run, routes
|
|
57
|
+
aio.py AsyncCronwatch, AsyncJobHandle, AsyncRunHandle: the client's methods as coroutines, the store in a worker thread
|
|
58
|
+
handler.py job.handler(): the SDK's fetch-style job handler, and its Django, Flask, Starlette, WSGI and ASGI adapters
|
|
59
|
+
_response.py which returned values are HTTP responses, and their status (a run failed by one of 400 or more)
|
|
60
|
+
celery.py Celery: the signal handlers, @cronwatch_task, install(), beat's schedules, the check task
|
|
61
|
+
apscheduler.py APScheduler 3: the listener, watch(), triggers as schedules
|
|
62
|
+
_convert.py what both share: cron fields from sets of values, "every" text, zone names, option checks
|
|
63
|
+
_scheduler_check.py a converted schedule checked against the scheduler's own fire times (the gem's check_fires)
|
|
64
|
+
web/
|
|
65
|
+
__init__.py Web (routes/index.ts): Request, Response, the routes, the WSGI app and the ASGI app
|
|
66
|
+
_html.py the pages (routes/html.ts), with the CSS text for text
|
|
67
|
+
_timeline.py the day and week timelines (routes/timeline.ts)
|
|
68
|
+
_pwa.py the manifest, app.js, sw.js and the icons' answers (routes/pwa.ts)
|
|
69
|
+
_icons.py the icons, written by scripts/make-dashboard-icons.mjs
|
|
70
|
+
_origin.py `new URL(value).origin`: the origin option and trust_proxy's forwarded origin
|
|
71
|
+
_escape.py escapeHtml, String(), truthiness, encodeURIComponent, toFixed, Object.entries
|
|
72
|
+
django/
|
|
73
|
+
__init__.py the settings (CRONWATCH), client(), routes(), reset(); DEBUG as the environment
|
|
74
|
+
apps.py the app config (label "cronwatch"); imports each app's cronwatch_jobs at startup
|
|
75
|
+
urls.py views.py the dashboard under the project's URLs
|
|
76
|
+
management/commands/cronwatch_check.py
|
|
77
|
+
stores/
|
|
78
|
+
__init__.py the Store protocol, MemoryStore, SqliteStore
|
|
79
|
+
memory.py
|
|
80
|
+
_sql.py sql.ts: the schema, statements, parameters and row mapping, text for text
|
|
81
|
+
sqlite.py
|
|
82
|
+
postgres.py PostgresStore on psycopg 3 (stores/postgres.ts)
|
|
83
|
+
tests/
|
|
84
|
+
test_conformance.py replays conformance/*.json (triage.json in test_triage.py)
|
|
85
|
+
test_client.py, test_client_hardening.py, test_start_finish.py the SDK's client tests
|
|
86
|
+
test_stores.py the store conformance test, memory, SQLite and Postgres
|
|
87
|
+
test_postgres.py the Postgres store's own tests (stores.test.ts), its schema, the app's transaction
|
|
88
|
+
test_channels.py the channels beside the replay: SigV4's AWS cases, SMS, refusals, urllib on local sockets
|
|
89
|
+
test_triage.py triage.json, and the real anthropic package's request on the wire
|
|
90
|
+
test_web_golden.py replays packages/ruby/test/web/golden.json: every page, header and JSON body of the SDK's routes
|
|
91
|
+
test_web.py the SDK's routes tests (routes*.test.ts), and the WSGI and ASGI apps
|
|
92
|
+
test_django.py cronwatch.django in a project made in the test (CI runs it on each Django series)
|
|
93
|
+
web_server.py the dashboard over wsgiref, for packages/mcp/test/python-web.test.ts
|
|
94
|
+
test_pgcron.py the pg_cron source on an in-memory cron schema and on a real pg_cron
|
|
95
|
+
test_node_compat.py one SQLite file shared with the built SDK (tests/node_store.mjs)
|
|
96
|
+
test_schedule_fuzz.py generated expressions answered by croner itself (tests/schedule_fuzz.mjs)
|
|
97
|
+
test_aio.py async jobs through JobHandle and AsyncCronwatch, the store kept off the event loop
|
|
98
|
+
test_handler.py handler() (the SDK's handler tests), its adapters, responses as outcomes
|
|
99
|
+
test_celery.py cronwatch.celery eagerly, on an in-process worker, on a prefork worker with Redis; beat and django-celery-beat
|
|
100
|
+
celery_app.py the Celery app that prefork worker runs
|
|
101
|
+
test_apscheduler.py cronwatch.apscheduler on background and asyncio schedulers, and events in every order
|
|
102
|
+
test_scheduler_check.py the schedule check against schedulers that agree, skip a run or run at other times
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
## The Python API
|
|
106
|
+
|
|
107
|
+
```python
|
|
108
|
+
import cronwatch
|
|
109
|
+
from cronwatch.stores import SqliteStore
|
|
110
|
+
|
|
111
|
+
cw = cronwatch.Cronwatch(
|
|
112
|
+
store=SqliteStore("./data/cronwatch.db"), # default: MemoryStore()
|
|
113
|
+
alerts=[cronwatch.Custom("pager", page)], # default: [Console()]
|
|
114
|
+
retention="30d",
|
|
115
|
+
)
|
|
116
|
+
# or cronwatch.configure(...) once, and cronwatch.client() wherever it is needed
|
|
117
|
+
|
|
118
|
+
nightly = cw.job("nightly-report",
|
|
119
|
+
schedule="0 2 * * *", timezone="UTC", grace="15m", timeout="30m",
|
|
120
|
+
expect="Report written", budget={"cost": 2}, failures_before_alert=1)
|
|
121
|
+
|
|
122
|
+
with nightly.run() as ctx: # a block
|
|
123
|
+
ctx.log("Report written:", path)
|
|
124
|
+
ctx.metric("cost", 1.2)
|
|
125
|
+
|
|
126
|
+
@nightly.monitor # a function: each call is a run
|
|
127
|
+
def build_report():
|
|
128
|
+
cronwatch.current().log("Report written")
|
|
129
|
+
|
|
130
|
+
nightly.run(lambda ctx: work(ctx)) # a function taking the context; returns what it returns
|
|
131
|
+
|
|
132
|
+
cw.check() # finds missed and stuck runs, retries undelivered alerts, prunes; a CheckResult
|
|
133
|
+
cw.start() # a daemon thread that checks every minute (long-running processes)
|
|
134
|
+
cw.silence("nightly-report", "2h")
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
`job(name, **options)` refuses an unknown option with `TypeError` and a bad value with `ValueError` (the SDK's messages). `run(fn)` and the decorator return what the function returns and raise what it raises, after the run is recorded; a returned string is the output when nothing was logged. `cronwatch.current()` is the run in progress in this thread (a `contextvars` variable, so it follows a task too). Options take durations as strings, milliseconds or `datetime.timedelta`, and `expect` as a string, a compiled `re.Pattern` (searched) or a function; a stored `timedelta` is written as its milliseconds and a pattern as JavaScript's `/source/flags`. `cw.run(name, fn, **options)` declares and runs. `defaults=` takes grace, timeout, timezone and failures_before_alert. The package ships `py.typed`: `job.run()` and `job.monitor` are overloaded so a type checker sees what they return (a function decorated with `@job.monitor` keeps its signature, `job.run(fn)` is what `fn` returns, an awaitable for an async one, and `job.run()` the block), and `job.handler()` is a `Handler`.
|
|
138
|
+
|
|
139
|
+
Sync is the base, since Celery tasks, Django management commands and cron scripts are sync. Async code is served two ways, both on the one synchronous orchestration:
|
|
140
|
+
|
|
141
|
+
- `JobHandle` takes async functions: `async with job.run() as ctx`, `@job.monitor` on an `async def` (the wrapper stays a coroutine function, each await a run) and `await job.run(async_fn)`. A run is split in two halves (`_Execution.record_start`/`record_end`, the store's, and `enter`/`leave`, the context's): the store's half runs in a worker thread (`asyncio.to_thread`), the function and `cronwatch.current()` in the caller's task, so concurrent runs each see their own context and the event loop never waits on the store, a channel or triage. A cancelled run is recorded as `Interrupted: CancelledError` and the cancellation goes on. A plain function that returns a coroutine (a lambda around an async call) is a failed run with a `TypeError`, never a quick success with the work left unawaited.
|
|
142
|
+
- `cronwatch.aio.AsyncCronwatch(client=None, **options)` is the client's methods as coroutines (`check`, `jobs`, `jobs_with_runs`, `job_summary`, `runs`, `get_run`, `silence`, `unsilence`, `forget`, `record_run`, `resume_run`, `close`), each the synchronous method in a worker thread; `job()` gives an `AsyncJobHandle` whose `run(fn)` is awaited for a plain function too (run in a worker thread), and whose `start()` and `resume()` give an `AsyncRunHandle` with `flush`, `finish` and `fail` awaited. `start()`, `stop()` and `routes()` are the synchronous client's, which never block. It wraps a synchronous client (`AsyncCronwatch(cronwatch.client())`), so sync and async code in one process share jobs, locks and state.
|
|
143
|
+
|
|
144
|
+
The plan was a second copy of the orchestration on async stores. It was not taken: the client's locks, compare-and-set loop, finish-once claim, channel threads and triage timeout would all have needed an async twin kept in step by hand, for no gain an app can see, since every store call already happens off the event loop. psycopg's async connection is not used for the same reason.
|
|
145
|
+
|
|
146
|
+
## Runs that span calls
|
|
147
|
+
|
|
148
|
+
`job.start(trigger=None, id=None)`, `job.resume(run_id)` and `cw.resume_run(name, run_id)` are the SDK's `start()`, `resume()` and `resumeRun()`, and return a `RunHandle`: `id`, `job`, `started_at`, `active`, `log`, `metric`, `metrics`, `flush()`, `finish(outcome=None, *, result=, error=)` and `fail(error)`. An outcome is None or `{"status": "ok"}` (ok), `{"error": e}` or `error=e` (failed, written like an error `run()` caught), or a string, `{"result": x}` or `result=x`, treated like `run()`'s return value, including the SDK's rule that an HTTP response of 400 or more fails the run (see Handler).
|
|
149
|
+
|
|
150
|
+
The client's side follows `client.ts`: a start with an id checks it (the SDK's messages, including the reserved `pgcron:` prefix) and holds a per-process lock keyed by the job and the id while it reads the store and inserts, so two starts with one id at once record one run, and another job's start with that id fails as it would one call later. `finish()` reads the stored run again, joins the stored output with the handle's (capped) and merges metrics (the handle's win), then judges it with the same code as `run()`. `flush()` redacts the lines it appends and writes only over a row still running and of this job (`update_run_if`); when it cannot, the handle keeps the lines for `finish()`, and it keeps the first 16 KB of what it flushed so `expect` at finish sees an early line. The store never raises out of `start`, `resume`, `flush` or `finish`: failures go to `on_error` as `recording <job>`, `starting <job>`, `resuming <job>`, `flushing <job>` or `finishing <job>`. A store that fails during `finish()` records nothing and leaves the handle active, lines kept, so it can be called again. A finish that records nothing (`... was already finished by this handle`, `... was already finished as ok`, `... was not found`, `... belongs to job "<other>"`) is reported, never raised, and returns None.
|
|
151
|
+
|
|
152
|
+
A run is judged once, however many processes finish it: the finish is written only over a stored row still `running`, else over one still `timeout` (a check already counted it as stuck: a late failure is written but not judged, a late success is judged and recovers), through the store's `update_run_if`, one conditional `UPDATE`. Only the process whose write lands evaluates. The stuck check marks a run timed out the same way, so a finish that landed meanwhile wins. `insert_run` raises for an id already stored, the memory store included.
|
|
153
|
+
|
|
154
|
+
## Threads and state
|
|
155
|
+
|
|
156
|
+
The client is synchronous; each run's work happens in the caller's thread. Every read-modify-write of a job's state goes through `Cronwatch._update_state`: holding the job's `RLock`, it reads the state, works out the next one, and writes it only when it changed, with `version` one higher, through the store's `compare_and_set_state` against the version read. A refused write is worked out again from a fresh read, up to 10 times. A store without `compare_and_set_state` gets `set_state`. Only store reads and writes happen under the lock; alerts are sent after it is released. Concurrent `check()` calls share one check (the first caller runs it, the others wait for its result). `start()` runs checks in a daemon thread named `cronwatch-check`, the first after a second and then on the interval; a second `start()` does nothing, and one with another interval is reported to `on_error`.
|
|
157
|
+
|
|
158
|
+
## Delivery
|
|
159
|
+
|
|
160
|
+
`deliver="now"` (the default) sends each alert from the process that produced it. `deliver="check"` sends nothing: the alert is queued in the job's state (`undelivered`, at most 20, the oldest dropped first and reported) for the next check in a process that delivers now, which triages and sends it, as the SDK's `deliver: "check"` does. An alert no channel accepted is queued the same way and retried once per check, oldest first; one that no longer describes the job (`evaluate.stale_alert`) is dropped; one check spends at most 20 seconds of wall clock retrying across all jobs. Triage is tried once per alert: `Alert.triage` None with `triage_tried` True is JSON `null`, never tried again.
|
|
161
|
+
|
|
162
|
+
Each channel sends in a thread of its own with a 15 second timeout, and triage gets 25 seconds. Python cannot stop a thread, so a channel (or triage) that times out is left to finish; until it has, nothing more is sent to it (the alert counts as not delivered there, and is retried) and no second triage starts, so a hung channel holds one thread rather than one per alert.
|
|
163
|
+
|
|
164
|
+
## Channels
|
|
165
|
+
|
|
166
|
+
A channel is an object with `name` and `send(alert, context)` that raises when the alert went nowhere; a plain function works too, and one that takes only the alert is called with the alert alone. `context.on_error(e)` reports a problem that did not stop the alert (one of several recipients refusing it) as `alert channel <name>`. `Console` (the default) and `Custom` are joined by the SDK's channels in `cronwatch.alerts`: `Slack`, `Discord`, `Webhook`, `Resend`, `Postmark`, `Sendgrid`, `Mailgun`, `Ses`, `Twilio`, `Sentry`, `Honeybadger`, `Datadog`, `Rollbar`, `Bugsnag` and `NewRelic`, keyword options in snake_case (`api_key`, `from_`, `subject_prefix`, `webhook_url`), constructor checks raising `ValueError` with those names (`Resend() needs an api_key`).
|
|
167
|
+
|
|
168
|
+
They are the SDK's, request for request, standard library only: the same URL, headers (lowercase names, in the SDK's order) and body, written with `_js.dumps` so the bytes are identical; the same stable alert id (the first 32 hex characters of SHA-256 of job, type and time) for idempotency keys, event ids and UUIDs; the same error text, `"<Provider> <origin> answered <status>: <body>"`, the body's secrets of four or more characters replaced by `[redacted]` in a prefix of 200 plus the longest secret's length, and only then cut to 200 UTF-16 units on a code point; only the URL's origin in any error; credentials trimmed before they go in a header, a webhook's header values too; lengths, cuts and SMS segments counted in UTF-16 units and GSM-7 septets as JavaScript counts them (a JSON body cut through a surrogate pair keeps the lone half, which `_js.dumps` escapes as `JSON.stringify` does, so even that byte matches); the same recovered defaults (sent by Sentry, Rollbar, Datadog, New Relic and the email channels; not by Twilio, Honeybadger or Bugsnag unless `recovered=True`). Each takes `http=` (anything with `post(url, body, headers)` returning a `Response(status, body)`), which is how the tests replay requests without a network. The default, `UrllibHTTP`, is `urllib.request` with its redirect handler refusing to follow (a 3xx comes back as an answer outside 2xx, so credential headers never go where it points), one ten second deadline for the whole request as `AbortSignal.timeout(10_000)` has it (each read bounded by what is left; past it before an answer it raises `RequestTimeout`, "The operation was aborted due to timeout"; past it while the body arrives it returns the status with an empty body, as the SDK reads a body it could not finish), header values trimmed as `fetch` trims them (one still holding a line break or NUL is refused with an error naming the header, never quoting its value, where `http.client`'s own error would quote it), URLs read as `fetch` reads them (spaces and control characters around them, and tabs and line breaks inside, dropped; only `http` and `https`, where urllib would also open `file:` and `ftp:`; a URL urllib cannot read is refused naming only its origin, never its path or query, which for a webhook is the credential), and bodies read as `response.text()` reads them (UTF-8, U+FFFD for bytes that are not, no byte order mark). Twilio sends to every number at once, a thread each, joined within the HTTP deadline plus a second; the alert is delivered when any number took it, each refusal reported through the context as `"<error> (to <mask>; <ok> of <n> numbers took the alert)"`, and it raises `"<first error> (<n> of <n> numbers failed)"` only when all refused. `tests/test_conformance.py` replays every case of `conformance/channels.json`: `sends` and `failures` for Slack, Discord and the webhook, `providerSends` and `providerFailures` for the other twelve (Twilio's parallel requests compared in the numbers' order), `twilioPartial`, and `textCuts`.
|
|
169
|
+
|
|
170
|
+
## Triage
|
|
171
|
+
|
|
172
|
+
`cronwatch.triage.anthropic.AnthropicTriage(context=..., model=..., effort=..., max_tokens=..., fallbacks=..., api_key=..., client=...)` is `triage/anthropic.ts`: the same default model (`claude-opus-5`), effort (`medium`), `max_tokens` (800), system prompt and user prompt text, byte for byte, the same `<job_data>` fencing of anything the job wrote, the same server-side fallback beta and `fallbacks: "default"` unless `fallbacks=False`, and one attempt (`with_options(max_retries=0)`) with a 24 second request timeout, under the client's 25 second wait. A field the installed anthropic package does not name yet goes in `extra_body`, so the body is the same whatever its version. A refusal, or an answer with no text, is None. `tests/test_triage.py` replays `conformance/triage.json` against a stub client and puts each request through the real package onto a mock transport, comparing the JSON on the wire.
|
|
173
|
+
|
|
174
|
+
## Sources
|
|
175
|
+
|
|
176
|
+
`sources=` takes objects with `name` and `sync(host)`, as the SDK's `sources` does. `check()` calls each, in order, after the store is ready and before anything else; one that raises is reported as `source <name>` and the check carries on, and the alerts a sync returns are added to the check's result. The host is the client: `job`, `record_run`, `store`, `now()` and `on_error(error, where)`.
|
|
177
|
+
|
|
178
|
+
`record_run(run, evaluate=True)` is the SDK's `recordRun`: keyed by the run's id, a new run is inserted, a stored run still `running` (or `timeout`, marked by a check) is finished once this one is not running, through the same claim as a handle's finish, and anything else is left alone. A stored run of another job is left alone and reported. Before that an `ok` run is checked against `expect`, and output and error are capped, redacted and cleared of NUL. The job must be declared in this process, or it raises `ValueError`.
|
|
179
|
+
|
|
180
|
+
`cronwatch.sources.pgcron.PgCron(db, jobs=, prefix=, job_name=, options=, timezone=)` is `sources/pgcron.ts` line for line, as the gem's `Sources::PgCron` is: the same SQL (settings read from `pg_settings`, which has no row for one the role may not read, so the caller's transaction is never aborted; the runs query a `LEFT JOIN` over `unnest` of the cursors, so open runs of any job are read), the same run ids (`pgcron:<prefix><runid>`), the same mapping from rows to runs (a finished row with no `start_time` starts at its `end_time`, else the job's newest run's start, else now; only the first five fields of a schedule are kept, `$` read as `L`), the per-job cursor found from the store the first time, the first-sight import of the twenty newest runs judged only from the newest finished one, a not-yet-started run held for `HOLD_MS` (ten minutes) then copied as running from when it was first seen, runs copied while going (or marked timeout since) read again until they finish, names no longer used for a job declared again without a schedule with the reason in the description, and the one-time warnings. `db` is anything with `query(sql, params)` returning rows as dicts, or a psycopg connection, a pool whose `connection()` lends one, or a connection string (a connection of its own, in autocommit mode). Statements go to Postgres with their `$n` placeholders through psycopg's `RawCursor`, lists as arrays, and timestamps come back as datetimes, read to the millisecond (text and epoch milliseconds are read too). A connection not in autocommit mode and outside a transaction is left as it was found (the reads are rolled back); inside the caller's transaction the reads join it and never end it. `conformance/pgcron.json` is replayed by `tests/test_conformance.py`, and `tests/test_pgcron.py` has the SDK's and the gem's tests, on an in-memory cron schema and, when `CRONWATCH_TEST_PGCRON` is set, on a real pg_cron.
|
|
181
|
+
|
|
182
|
+
## Keeping in step
|
|
183
|
+
|
|
184
|
+
`conformance/` at the repo root holds JSON cases generated from the TypeScript build by `scripts/conformance.mjs`. `tests/test_conformance.py` replays every case (duration, schedule, evaluate, format, health, output, the store scripts against the memory, SQLite and Postgres stores, the channels and the pg_cron source; `triage.json` in `tests/test_triage.py`), comparing values as the JSON the SDK writes, and a new fixture file fails the suite until it is replayed. A behaviour change lands in TypeScript first, `npm run conformance` regenerates the fixtures, and this package is fixed until they pass. The dashboard is held to the SDK the same way, by the routes fixture the gem replays too (see Web).
|
|
185
|
+
|
|
186
|
+
Croner parity is also checked against croner itself: `tests/test_schedule_fuzz.py` generates 3,000 expressions (valid and malformed, nicknames, names, ranges, steps, lists, `L`, `W`, `LW`, `#`, `?`, `+`, six fields) in zones with and without daylight saving, from times around the clock changes, and the SDK in Node must give the same error message or the same fire times. Unlike the gem, the port reads every form croner reads (`W`, `LW`, `5L` in the day of the month, a year field, a range with `#`), because it ports croner's parser rather than wrapping another.
|
|
187
|
+
|
|
188
|
+
## Storage
|
|
189
|
+
|
|
190
|
+
`SqliteStore` writes the same three tables as `packages/sdk/src/stores/sql.ts`: same names (`cronwatch_` prefix by default, the same prefix rules), same columns and types, the same `CREATE` text (so `sqlite_master` reads the same whoever created them), the same statements, and the same JSON in the JSON columns. WAL mode, `busy_timeout` 5000 and `synchronous` NORMAL, the journal mode switched with the SDK's retry of a busy database, its directory created when missing and the file with mode 0600. One connection per store, shared by threads under a lock, in autocommit mode; `delete_job` is one transaction. `tests/test_node_compat.py` has the built SDK and this store replay the same store calls into two files and compares what each reads of the other's and every column's bytes and SQLite type, and has a Node client and a Python client take turns on one file and on one job's state version.
|
|
191
|
+
|
|
192
|
+
`cronwatch.stores.postgres.PostgresStore(conninfo=None, pool=None, prefix=...)` is `stores/postgres.ts` on psycopg 3: the same schema (`BIGINT` times, `JSONB`, `BIGSERIAL seq` breaking ties between runs started in one millisecond, names sorted `COLLATE "C"`), the same statements with `$n` placeholders sent as written through psycopg's `RawCursor`, the same parameters and JSON, `init` under `pg_advisory_xact_lock(hashtext('cronwatch:<prefix>'))` so many processes can start at once, `delete_job` in one transaction, and `update_run_if` and `compare_and_set_state` as single conditional statements. It reads `DATABASE_URL` when given no connection string. It writes through a connection of its own, in autocommit mode, shared by threads under a lock, opened again when it breaks and left to the parent in a forked child, so its writes never join a transaction the app has open (a run recorded inside one survives a rollback). `pool=` takes the app's psycopg_pool pool instead (anything whose `connection()` lends a connection), which is not closed by `close()`. It passes the store conformance test, `store.json`'s scripts, and the finish-once tests with several stores over one database, as the memory and SQLite stores do, when `CRONWATCH_TEST_PG` is set.
|
|
193
|
+
|
|
194
|
+
## Web
|
|
195
|
+
|
|
196
|
+
`cw.routes(token=, base_path=, origin=, trust_proxy=)` is the SDK's `cw.routes()`: a `cronwatch.web.Web` whose `handle(Request) -> Response` is `routes/index.ts` line for line (the same URLs, JSON, status codes, headers, CSP, cookie, redirects, cross-site rule, token and cron secret rules, development token and its sign-in line, and 500s reported to `on_error` as `routes`), and whose pages are `routes/html.ts` and `routes/timeline.ts` byte for byte, the CSS included. The `Web` is a WSGI app, and `Web.asgi` is the same routes as an ASGI app (the lifespan answered, each request handled in a worker thread with `asyncio.to_thread`, since the client underneath is synchronous, the async client's included). `tests/test_web_golden.py` replays the gem's fixture, `packages/ruby/test/web/golden.json` (57 captures of the SDK's routes over a fixed seed, PNGs as base64), through the WSGI app and through `handle`, and every status, header and body matches; `tests/test_web.py` has the SDK's routes tests. A behaviour change lands in TypeScript first, and `TZ=UTC node packages/ruby/test/web/golden.mjs` regenerates the fixture for both ports.
|
|
197
|
+
|
|
198
|
+
What a Python server hands over is read the way the SDK reads a fetch `Request`: the path as the browser sent it (still percent-encoded, mount point included), headers as latin-1 text, the query as `URLSearchParams` parses it, the body as `request.json()` and `request.formData()` read it (JSON values as `String(value)` gives them, a file field as `[object File]`), and the origin of the request's own URL, lowercased and without a default port. A server decodes the path before the app sees it, so the raw request target is used when the server passes one that decodes to the same path (gunicorn's `RAW_URI`, uWSGI's `REQUEST_URI`, ASGI's `raw_path`), and otherwise the decoded path is encoded again as a browser encodes it. `trust_proxy` follows the SDK (off by default; when on, the first `X-Forwarded-Proto` and `X-Forwarded-Host`, ignored unless they make a scheme and a bare host). The base path defaults to the mount point (`SCRIPT_NAME`, ASGI's `root_path`), as the gem's does, rather than the SDK's `"/cronwatch"`, since a WSGI or ASGI mount says where it is. A `HEAD` answer has no body (and the `Content-Length` a `GET` would have).
|
|
199
|
+
|
|
200
|
+
A request body is read into memory up to `web.MAX_BODY` (1 MiB; the routes' forms and JSON are a few bytes). An ASGI server hands the body to the app before the routes can ask for a token, so the ASGI apps (the dashboard's and a handler's) stop reading once a body passes it and answer 413 `{"ok":false,"error":"Request body too large"}`; the WSGI app reads a body only once a route wants it, after the token, and answers the same to a `Content-Length` over the limit (a negative one reads nothing). The SDK leaves this to the server in front of it.
|
|
201
|
+
|
|
202
|
+
## Django
|
|
203
|
+
|
|
204
|
+
`cronwatch.django` is an app (`INSTALLED_APPS += ["cronwatch.django"]`, label `cronwatch`) and a URLconf (`path("cronwatch/", include("cronwatch.django.urls"))`): every path under the prefix goes to one view, which reads the Django request into the routes' `Request` (the base path is where the URLs were included; the origin is `request.scheme` and `request.get_host()`, so `SECURE_PROXY_SSL_HEADER` and `USE_X_FORWARDED_HOST` apply; the raw path from `RAW_URI`, `REQUEST_URI` or the ASGI scope; a form Django already read from the stream from `request.POST`) and writes the `Response` back as it is, `Set-Cookie` included as the routes' own header text. The view is `csrf_exempt`: the forms carry no Django token, and the routes refuse a cross-site write themselves, as the SDK does.
|
|
205
|
+
|
|
206
|
+
Settings are one dict, `CRONWATCH`: the client's options in upper case (`STORE`, `ALERTS`, `TRIAGE`, `SOURCES`, `CRON_SECRET`, `RETENTION`, `DEFAULTS`, `REDACT`, `DELIVER`, `ON_ERROR`; a store, channel or source may be a dotted path, to the thing or to a class or function that makes it), or `CLIENT` for a client the app made itself, and the dashboard's (`TOKEN`, `BASE_PATH`, `ORIGIN`, `TRUST_PROXY`); an unknown key raises `ImproperlyConfigured`. `cronwatch.django.client()` makes the client from them on first use with `cronwatch.configure`, so `cronwatch.client()` is the same client; with no client options it is `cronwatch.client()`. A change to `CRONWATCH` or `DEBUG` (`override_settings` in a test) drops the client and dashboard made from them. `python manage.py cronwatch_check` runs `client().check()` and prints `cronwatch: checked N jobs, sent M alerts` (nothing at `--verbosity 0`), as the gem's `cronwatch:check` task does; it checks every job in the store, so a job declared in a module the command never imports is still checked from its stored definition. `DEBUG` is the environment when no variable names one (see Rules).
|
|
207
|
+
|
|
208
|
+
Jobs are declared where Django can find them at startup, as the gem's Railtie declares every monitored job at boot: the app config's `ready()` imports a `cronwatch_jobs` module from each installed app that has one (`autodiscover_modules`, as the admin finds `admin.py`), so the jobs declared there (`client().job(...)`) are known before their first run, and `cronwatch_check` reports one that never ran. A module that fails to import stops the startup, as a bad declaration stops the gem's boot. A test that overrides `CRONWATCH` gets a new client, without the jobs declared on the one before. Django 5.2 stays the oldest series supported (4.2 is out of support).
|
|
209
|
+
|
|
210
|
+
`tests/test_django.py` makes a project in the test (the dashboard under a two-part prefix, Django's security, common, CSRF and clickjacking middleware, and two job handlers as views) and checks the pages through Django against `handle`'s byte for byte, the forms without a Django CSRF token, the sign-in cookie, `DEBUG`, the command, the settings, Django's WSGI handler and its ASGI side, the handler views (exempt from CSRF, plain and async), and, in a project and process of its own, the startup import of `cronwatch_jobs`. It runs on the dev group's Django; CI runs it on 5.2 LTS (Python 3.11 and 3.14), 6.0 and 6.1 (Python 3.14; Django 6 needs 3.12 or newer), the series supported as of this release, with `uv run --with "django~=X.Y.0"`.
|
|
211
|
+
|
|
212
|
+
## Handler
|
|
213
|
+
|
|
214
|
+
`job.handler(fn, secret=...)` is the SDK's `handler()`, for a platform cron that calls a URL: `fn(ctx, request)` runs as a recorded run with the trigger `"handler"` for each request whose `Authorization` is `Bearer <secret>` (compared in constant time). The secret is the handler's own, else the client's `cron_secret` (`$CRON_SECRET` by default, "" counting as unset); `secret=None`, or a client made with `cron_secret=None`, lets anyone run the job. With no secret and no opt-out, outside development, it answers 503 with the SDK's JSON and reports to `on_error` once as `handler`; a wrong or missing bearer is 401 `{"ok":false,"error":"Unauthorized"}`. A run is answered with `{"ok","job","run","status","durationMs"}`, 200 or 500, and the error's first line as `"error"` only when a secret was required; the same bytes and headers (`content-type: application/json; charset=utf-8`, `cache-control: no-store`) as the SDK's `json()`. A function that returns a response is answered with it as it is. An exception is answered with the 500; an interrupt (`KeyboardInterrupt`, `CancelledError`) is recorded and raised.
|
|
215
|
+
|
|
216
|
+
A response, returned by any run's function or given to `finish(result=)`, fails the run when its status is 400 or more, with the error `HTTP <status>` and its reason phrase when it carries one (`HTTP 503 Service Unavailable`), as the SDK has a fetch `Response`. Python has no one response type, so `_response.py` reads the common ones by duck typing: `cronwatch.web.Response`, anything with a whole-number `status_code` from 100 to 599 (Django, Flask and Werkzeug, Starlette and FastAPI, requests, httpx; the reason from `reason_phrase`, `reason` or Werkzeug's status line), and `http.client.HTTPResponse`.
|
|
217
|
+
|
|
218
|
+
A request is anything with headers: a `cronwatch.web.Request`, a framework's request (headers read without regard to case), a WSGI environ, or a serverless event with a `headers` mapping. Calling the handler answers in the request's own kind (a Django `HttpResponse`, a Werkzeug response, a Starlette response, else a `cronwatch.web.Response`), and `.django`, `.flask`, `.starlette`, `.wsgi` and `.asgi` are views, endpoints and apps for each (Django's exempt from CSRF: the caller proves itself with the bearer, not a cookie). An async function gives an `AsyncHandler`, awaited, whose adapters are async too; its WSGI app runs it on an event loop of its own.
|
|
219
|
+
|
|
220
|
+
`.aws_lambda(event, context)` is the handler as an AWS Lambda function (`lambda_handler = cron.aws_lambda`) behind API Gateway or a function URL: the bearer is read from the event's `headers` without regard to case (a REST API's are as sent, an HTTP API's and a function URL's lowercased), or from a REST API's `multiValueHeaders` when `headers` has none; the function gets the event; and the answer is the proxy result, `{"statusCode", "headers", "body", "isBase64Encoded"}`, the body as text when it is UTF-8 and base64 otherwise, `Content-Length` left out (API Gateway sets its own) and a header given twice joined with `, `. No other key is written, since a REST API refuses a proxy result with keys it does not know. An async function's handler runs it to completion with `asyncio.run`, as the Lambda runtime calls a plain function. A Lambda proxy result is a response too (`_response.py`: a dict with a whole-number `statusCode` from 100 to 599 and none but the proxy result's keys), so a function that returns one is answered with it, and one of 400 or more fails the run, from any adapter and from `run()`. A directly invoked function (EventBridge Scheduler) has no headers to carry a bearer; IAM decides who may invoke it, so such a handler takes `secret=None`.
|
|
221
|
+
|
|
222
|
+
`tests/test_handler.py` has the SDK's handler tests and the adapters against Flask, Starlette, Werkzeug, raw WSGI and ASGI, and REST API, HTTP API and function URL events; `tests/test_django.py` the Django views.
|
|
223
|
+
|
|
224
|
+
## Celery
|
|
225
|
+
|
|
226
|
+
`cronwatch.celery.install(app, client=, beat=True, django_celery_beat=None, tasks=, exclude=, **options)` watches the app's tasks through Celery's signals, so no task changes: `task_prerun` starts a run (the `_Execution` of `run()`, with the run id `<task id>:<12 hex>`, and `cronwatch.current()` its context), and `task_success`, `task_failure` and `task_retry` end it, in the worker process that ran the task; `task_postrun` ends one they did not (Celery's `Ignore` is an ok run, `Reject` a failed one, and an exception past Celery's own handling, an interrupt or an eager task with `task_eager_propagates`, is read from the exception in flight). The run's state rides on the task's request between the signals, so nested and concurrent tasks (the threads pool, gevent) keep theirs apart. A task that raises still raises on to Celery, so its retries, error handlers and result backend see it unchanged.
|
|
227
|
+
|
|
228
|
+
Which tasks: every task beat schedules, and any with `@cronwatch_task(**options)` (below `@app.task` on the function, or above it on the task; a task not made yet is read once it is) or an entry in `install(tasks={name: options})`, named after the task unless `name=` says otherwise. `exclude` leaves out beat entries by key and tasks by name; `celery.backend_cleanup` and the check task are never jobs. install's own options apply to every job it declares; a task's own win, and a task's own `schedule` replaces beat's.
|
|
229
|
+
|
|
230
|
+
Schedules come from `app.conf.beat_schedule` and, when django-celery-beat is installed (or `django_celery_beat=True`), its enabled recurring `PeriodicTask` rows (clocked and one-off rows are not schedules), which win over `beat_schedule` entries of the same name, as its scheduler copies those into the table. A crontab becomes the five fields croner reads, written from the sets of values Celery matches (`*`, `*/n` for three or more, lists with ranges), in the crontab's zone (the app's `timezone`, UTC by default, the process's IANA zone when `enable_utc` is off, or django-celery-beat's per-row zone); Celery fires on a day matching the day of the month, the month and the day of the week together, which croner reads with `+` before the day of the week when both days are restricted. An interval becomes `every <interval>`. Each converted crontab is checked against Celery's own `crontab.is_due`, asked as beat asks it, with `_scheduler_check.py` (the gem's `check_fires`: around every clock change in the next five years and from the start of each month of a sample year; the answer cached per expression and zone). A schedule that cannot be converted (a solar or other schedule, a task scheduled by several entries, one never due, one CronWatch would expect at other times) is reported once per process to `on_error` as `declaring <entry>`, and the task is watched without a schedule.
|
|
231
|
+
|
|
232
|
+
Declared when: at `worker_init` (in the worker's main process, before the pool forks), again at the first run of a decorated task not seen yet, whenever a minute has passed since beat's schedule was read, and by the check task. `cronwatch.celery.check`, a shared task, declares every watched job and runs `client.check()`; schedule it with beat, once for the deployment. `watch.declare()` raises for the first problem, for a startup that should stop on one.
|
|
233
|
+
|
|
234
|
+
Retries follow the gem's Sidekiq rules: every attempt is a run of its own (a retry keeps the task id, so its runs share the id's prefix), and an attempt that ends in a retry is a failed run with the error that caused it (`Retry.exc`), so failing attempts open one failed alert and the one that succeeds closes it; `failures_before_alert` rides through a number of them. A run whose pool process is gone is failed by the worker's main process when Celery reports it there: `WorkerLostError` and a hard time limit (`task_failure`), a revoke with terminate (`task_revoked`), a task cancelled when the broker connection dropped (`task_retry` with no reason); the main process finds the running runs by the task id's prefix. One Celery never reports (a requeued task after the worker died) is marked stuck by a check after the job's timeout.
|
|
235
|
+
|
|
236
|
+
The client is `client=` (a client or a function returning one), else the Django integration's in a Django project that uses it, else `cronwatch.client()`, looked up when used. A pool process spawned rather than forked (macOS) has the app made again from a pickle, so an app of the same `main` name that `install()` was given is taken for it. `SqliteStore` opens a connection of its own in a forked child (the parent's is left to it), as `PostgresStore` does, so the prefork pool can share a file.
|
|
237
|
+
|
|
238
|
+
`tests/test_celery.py` runs tasks eagerly (`apply`, `task_always_eager`), on a real worker in the test process (the in-memory broker, the solo and threads pools, through `celery.contrib.testing`), and, when `CRONWATCH_TEST_REDIS` is set, on a real prefork worker (`tests/celery_app.py`) whose children record to one SQLite file, where a hard time limit and a child that exits are failed from the main process. django-celery-beat's table is read in a process of its own with its own settings. CI runs them on Celery 5.6 with a Redis service; Celery 5.5 is the oldest supported.
|
|
239
|
+
|
|
240
|
+
## APScheduler
|
|
241
|
+
|
|
242
|
+
`cronwatch.apscheduler.watch(scheduler, client=, jobs=, exclude=, **options)` adds a listener to an APScheduler 3 scheduler (background, blocking, asyncio or any other). Every job is declared, named after its id (or, for an id APScheduler generated, 32 hex characters, its name, which is the function's), with its trigger as the schedule: a `CronTrigger` becomes the cron expression croner reads in the trigger's zone (Monday is APScheduler's 0, `last` is `L`, `1st fri` is `5#1`, `last fri` is `5L`, both days restricted is `+`, seconds a sixth field), checked against the trigger's own `get_next_fire_time` as Celery's crontabs are; an `IntervalTrigger` is `every <interval>`; a `DateTrigger` runs once and has no schedule; anything else (combined triggers, calendar intervals, a week or year field, jitter aside) is reported and the job watched without one. A job added or rescheduled is declared again, and a removed one is declared again without its schedule, so it is not reported missed.
|
|
243
|
+
|
|
244
|
+
A run starts when APScheduler reports the job submitted to its executor and ends with its executed or error event: the return value is the output when it is a string (and checked by `expect`, and failed by an HTTP response of 400 or more), an exception fails it. A run skipped because it could not start within `misfire_grace_time` is failed with that reason. The events are recorded in order in a thread of the listener's own (`cronwatch-apscheduler`), so neither the scheduler's thread nor an asyncio scheduler's loop waits on the store; `flush()` waits for them, `close()` stops listening. A run that ends before its submission is heard of (a fast job: APScheduler dispatches the submission after its loop over due jobs) is recorded once, as it ends. Catch-up runs of one submission (`coalesce=False`) are reported together when the last ends, so each after the first is recorded when it is reported, with no duration.
|
|
245
|
+
|
|
246
|
+
APScheduler 4 is not supported: it is still a pre-release (4.0.0a6), and it replaced the events, triggers and scheduler API this is built on. `tests/test_apscheduler.py` covers the conversions, a background and an asyncio scheduler, and events dispatched by hand in the orders a busy scheduler delivers them.
|
|
247
|
+
|
|
248
|
+
## Phases
|
|
249
|
+
|
|
250
|
+
1. Done: the design, the core (duration, schedule, evaluate, stats, output, format, serialize, the client with runs, handles, sources, record_run, deferred delivery and triage hooks), the memory and SQLite stores, conformance and the SDK's client tests.
|
|
251
|
+
2. Done: stores and integrations that need no web: Postgres (psycopg 3), the alert channels, Claude triage (`anthropic`), the pg_cron source.
|
|
252
|
+
3. Done: the web dashboard and JSON API (WSGI and ASGI), matching the SDK's routes byte for byte, and the Django integration (settings, a management command `cronwatch_check`, the app's URLs, `DEBUG` as the environment).
|
|
253
|
+
4. Done: Celery (signal handlers, a `@cronwatch_task` decorator, beat's and django-celery-beat's schedules read as job schedules, the check task), APScheduler 3 (a listener recording each job's runs, its triggers as schedules), async jobs and the async client, `handler()`, and the Django app importing each app's `cronwatch_jobs`.
|
|
254
|
+
|
|
255
|
+
## Where it cannot match the SDK
|
|
256
|
+
|
|
257
|
+
- A cron expression that names a date no month has (`0 0 30 2 *`) makes croner, which walks by recursion a year at a time, run out of stack before the year 3000, so the SDK reports the job as unevaluable. The port walks in a loop and answers that the schedule never fires: no next expected time, never missed.
|
|
258
|
+
- Croner reads a string with a colon after its first character as a one-time date, through JavaScript's lenient `Date.parse`, so the SDK accepts `2026-12-01T00:00:00` (or even `0 2:30 * * *`) as a schedule that fires once, or never. The port refuses every such string: one that looks like an ISO date with `CronPattern: a one-time date is not supported by the Python port`, anything else with the message croner gives for text `Date.parse` cannot read (`Invalid ISO8601 passed to timezone parser.`).
|
|
259
|
+
- In the process's own zone (no `timezone`), a wall-clock time in a gap is found by croner's `fromTZ` rule, as it is for a named zone, where JavaScript's local `Date` uses the offset before the transition. The two agree for one-hour gaps; the conformance fixtures run in UTC.
|
|
260
|
+
- `handler()`'s refusal without a secret names the Python API, `pass secret=None to handler()` (and `pass secret=None to allow unauthenticated requests` in the report to `on_error`), where the SDK says `secret: null`. A response is any of the kinds `_response.py` reads rather than one fetch `Response`, and its reason phrase is the one the object carries (Django fills one in for every status, so a Django 503 fails a run with `HTTP 503 Service Unavailable` where a fetch `Response` made without a `statusText` gives `HTTP 503`).
|
|
261
|
+
- The async client is the synchronous one with its store calls, alerts and triage in worker threads, not a second orchestration on async stores (see The Python API), so it needs asyncio (it uses `asyncio.to_thread`), as the ASGI apps do; trio and curio are not supported.
|
|
262
|
+
- When the server passes only a decoded path (wsgiref, Django's `runserver`) and no raw request target, an encoded `/` in a path segment (`%2F`) reads as a separator and a malformed escape (`%zz`) as a literal `%`, so such a request answers 404 where the SDK answers 400. Job names made by `cw.job` hold neither character; only a job another writer stored could. The SDK also reads the path after the URL parser has resolved `.` and `..` segments; the port takes it as sent, which is what a browser sends.
|
|
263
|
+
- Two messages name the Python API rather than the SDK's: the locked page (`pass token= to cw.routes(), or TOKEN in Django's CRONWATCH setting, or pass token=None`) and the empty board (`cw.job("name", schedule="0 2 * * *")`), as the gem names its own.
|
|
264
|
+
- A non-ASCII host in the `origin` option becomes punycode through Python's `idna` codec (IDNA 2003), where the URL parser uses UTS 46: they differ for a few characters (`ß`, final sigma, joiners). Write such a host in punycode to be exact.
|
|
265
|
+
- The ASGI app runs each request in a worker thread with `asyncio.to_thread`, so it needs an asyncio server (uvicorn, Hypercorn on asyncio, Daphne), not trio.
|
|
266
|
+
- The SDK has no Celery or APScheduler integration; these follow the gem's Sidekiq and scheduler integrations. A Celery task scheduled by several beat entries has one job and no schedule (the gem refuses such a class too): the worker cannot tell which entry sent a run. A job's schedule is read from beat's configuration, so beat's own runtime state (`beat_cron_starting_deadline` skipping a run, a schedule changed in beat's shelve file) is not seen.
|
|
267
|
+
- APScheduler tells a listener only that a job was submitted and how it ended, so inside an APScheduler job `cronwatch.current()` is None (the return value is the output instead), a run starts when the job is submitted rather than when its thread picks it up, and catch-up runs of one submission after the first have no duration. APScheduler 3 answers the fire time after a cron trigger's second occurrence of a repeated wall time (the second 01:30 when clocks go back) with that same time, a fault of APScheduler's own; the schedule check steps past it rather than take it as a mismatch.
|
|
268
|
+
- Under Django, the framework's own middleware may add headers the SDK does not send (`Cross-Origin-Opener-Policy` from `SecurityMiddleware`, `X-Frame-Options` on JSON from `XFrameOptionsMiddleware`); the routes' own headers and bodies are unchanged. With `DEBUG` off and no variable naming the environment, the environment is production, so a Django test run (which turns `DEBUG` off) gets the in-memory store's warning and a dashboard without a token answering 503, as a production process would.
|
|
269
|
+
- A channel whose request is redirected fails with `"<Provider> <origin> answered 307"` (or Slack's, Discord's and the webhook's own `answered 307`), since urllib is told not to follow and hands the 3xx back; the SDK asks `fetch` for `redirect: "error"`, which fails with a network error of its own instead. Either way nothing is sent to where the redirect points and the alert counts as not delivered. The gem is the same.
|
|
270
|
+
- `urllib.request` honours the `HTTP_PROXY`, `HTTPS_PROXY` and `NO_PROXY` environment variables, where Node's `fetch` ignores them unless told otherwise, so a Python process behind a proxy sends its alerts through it.
|
|
271
|
+
- The anthropic package cannot abandon a request under way when the client stops waiting, so triage checks the abort signal before the request starts and relies on the 24 second request timeout after that (the SDK passes the signal to the request). The client already moves on after 25 seconds either way. The run text triage quotes, cut through a surrogate pair, ends in U+FFFD where the SDK's request would carry the lone half, because the anthropic package writes its JSON as UTF-8, which cannot hold one.
|
|
272
|
+
- The pg_cron source reads a run id back from the store only when what follows its prefix is 1 to 15 digits, as the gem does; the SDK's `Number()` would also read `pgcron:` alone as 0, or `pgcron:1e3` as 1000. The source never writes such ids.
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Jon Phillips
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|