cnpg-drill 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 cnpg-drill contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,131 @@
1
+ Metadata-Version: 2.4
2
+ Name: cnpg-drill
3
+ Version: 0.1.2
4
+ Summary: Prove CloudNativePG Barman backups restore into an isolated cluster
5
+ Author: cnpg-drill contributors
6
+ License: MIT License
7
+
8
+ Copyright (c) 2026 cnpg-drill contributors
9
+
10
+ Permission is hereby granted, free of charge, to any person obtaining a copy
11
+ of this software and associated documentation files (the "Software"), to deal
12
+ in the Software without restriction, including without limitation the rights
13
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
14
+ copies of the Software, and to permit persons to whom the Software is
15
+ furnished to do so, subject to the following conditions:
16
+
17
+ The above copyright notice and this permission notice shall be included in all
18
+ copies or substantial portions of the Software.
19
+
20
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
21
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
22
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
23
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
24
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
25
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
26
+ SOFTWARE.
27
+
28
+ Keywords: cloudnativepg,postgresql,kubernetes,backup,restore
29
+ Classifier: Programming Language :: Python :: 3
30
+ Classifier: License :: OSI Approved :: MIT License
31
+ Classifier: Operating System :: OS Independent
32
+ Requires-Python: >=3.10
33
+ Description-Content-Type: text/markdown
34
+ License-File: LICENSE
35
+ Dynamic: license-file
36
+
37
+ # cnpg-drill
38
+
39
+ **Prove that a CloudNativePG backup can restore.** `cnpg-drill` creates a disposable CloudNativePG cluster from the source cluster's Barman Cloud object store, waits for PostgreSQL to become ready, runs read-only SQL assertions, emits a JSON result, and deletes the drill cluster. It can run locally or on a schedule as a Kubernetes CronJob.
40
+
41
+ The core is free and works without an account or external service. It does not make backups or copy database contents to a vendor. The proposed hosted service in [RESEARCH.md](RESEARCH.md) is separate and has not been built.
42
+
43
+ ## Status
44
+
45
+ **Early alpha.** Unit tests cover manifest safety and execution state. A disposable [local integration run](integration/local/README.md) on v0.1.0 proved full restore, PITR, read-only archive access, denied-WAL reporting, failed-check reporting, and Cluster/PVC cleanup on one version matrix. The v0.1.1 release archive and public chart 0.1.2 also restored a fresh backup and checked a row in an application database. A separate single-node pgEdge chart restore passed SQL checks in its `app` database. Validate against your operator, PostgreSQL image, storage, object store, and Barman plugin versions before treating a pass as disaster recovery assurance.
46
+
47
+ The [v0.1.2 release](https://github.com/danielgaskins/cnpg-drill/releases/tag/v0.1.2) carries literal AWS checksum compatibility settings into recovery clusters without copying source credentials. It includes the MIT-licensed kubectl plugin archive and checksum. The [default Krew index submission](https://github.com/kubernetes-sigs/krew-index/pull/6373) and [custom Krew index submission](https://github.com/ishantanu/awesome-kubectl-plugins/pull/44) currently package v0.1.0 and are under review. [kubetools](https://github.com/collabnix/kubetools/pull/431) is reviewing a Backup Tools listing.
48
+
49
+ The [Helm chart is listed on Artifact Hub](https://artifacthub.io/packages/helm/cnpg-drill/cnpg-drill) as a Verified Publisher package. Chart version `0.1.2` uses the v0.1.1 application image and includes the publisher's [website](https://danielgaskins.com/).
50
+
51
+ ## What it supports
52
+
53
+ - One source CloudNativePG cluster with one enabled Barman Cloud plugin and a named `ObjectStore` in the same namespace.
54
+ - An optional separate `recoveryObjectStore` with credentials limited to archive reads. Preflight checks that its destination and endpoint match the source store before creating a Cluster.
55
+ - Full restore from the latest completed Barman plugin `Backup` resource, or PITR to an explicit RFC3339 `targetTime`. The selected backup ID is pinned and included in the report; a stale backup fails preflight.
56
+ - Single-instance drill cluster using the source's PostgreSQL image, storage size/class, optional WAL storage, and resource requests.
57
+ - Copies literal AWS checksum compatibility settings from the source Cluster for S3-compatible archive restores; it does not copy writer credentials or arbitrary environment variables.
58
+ - Read-only SQL checks in `postgres` or a named application database. Exact expected values make reports easy to audit.
59
+ - Optional timestamp freshness assertions (`maxAgeSeconds`) for an application-defined recovery point. This measures age of that application's data, not the database WAL RPO.
60
+ - A JSON report on stdout and optionally in a local file. Query outputs are hashed rather than logged. Exit code 0 means all checks, Cluster deletion, and PVC cleanup passed; 1 means the drill failed; 2 means invalid input or a failed preflight.
61
+ - Scheduled execution with the Helm CronJob in `deploy/helm/cnpg-drill`.
62
+
63
+ The CLI refuses source clusters with tablespaces or recovery bootstrap because a naive clone could produce an incomplete or misleading pass. Volume snapshots, cross-namespace restores, cross-region recovery, custom Postgres extension image changes, and multi-cluster fleet management are future work.
64
+
65
+ ## Install and use
66
+
67
+ For a first run on your own cluster, follow the [first-run guide](docs/FIRST-RUN.md).
68
+
69
+ Requirements: Python 3.10+, `kubectl` in `PATH`, CloudNativePG and the Barman Cloud plugin installed in the target cluster, and Kubernetes access to read the source Cluster and ObjectStore, create/get/delete a drill Cluster, and exec into its PostgreSQL pod.
70
+
71
+ ```bash
72
+ python3 -m pip install -e .
73
+ cnpg-drill plan --config examples/drill.json
74
+ cnpg-drill run --config examples/drill.json --report reports/latest.json
75
+ ```
76
+
77
+ `plan` reads the source configuration and prints the exact Cluster manifest it would create. Review that manifest before running the first drill. The example's table-count assertion is illustrative; replace it with an invariant from your own application. A minimal check is `SELECT 1`, but it only proves connection to the recovered server, not useful application data.
78
+
79
+ Example `drill.json`:
80
+
81
+ ```json
82
+ {
83
+ "namespace": "production",
84
+ "cluster": "app-db",
85
+ "recoveryObjectStore": "app-db-recovery-readonly",
86
+ "timeoutSeconds": 1800,
87
+ "maxBackupAgeSeconds": 691200,
88
+ "checks": [
89
+ {"name": "postgres-ready", "query": "SELECT 1", "expected": "1"},
90
+ {"name": "orders-exist", "database": "app", "query": "SELECT count(*) > 0 FROM public.orders", "expected": "t"},
91
+ {"name": "recent-order", "database": "app", "query": "SELECT max(created_at)::timestamptz FROM public.orders", "maxAgeSeconds": 86400}
92
+ ]
93
+ }
94
+ ```
95
+
96
+ For PITR, add `"targetTime": "2026-09-29T10:00:00Z"`. Pick a time inside the backup and WAL retention window. A passing PITR check proves only that target; use a second drill for the latest recovery path.
97
+
98
+ ## Scheduled drill
99
+
100
+ The Helm chart creates a namespace-scoped ServiceAccount, Role, RoleBinding, ConfigMap, and CronJob. Set `cluster`, `checks`, and the other settings in a values file and install it **in the same namespace as the source Cluster and ObjectStore**. The chart defaults to suspension so installation does not immediately start a restore.
101
+
102
+ The example image build pins kubectl 1.36.4, suitable for Kubernetes 1.35–1.37 under the project's [version skew policy](https://kubernetes.io/releases/). Set `KUBECTL_VERSION` at image build time for another supported cluster version.
103
+
104
+ ```bash
105
+ helm install cnpg-drill oci://ghcr.io/danielgaskins/charts/cnpg-drill \
106
+ --version 0.1.2 -n production -f drill-values.yaml
107
+ ```
108
+
109
+ See the [first-run guide](docs/FIRST-RUN.md) and [chart values](deploy/helm/cnpg-drill/README.md) for a suspended first run with a separate read-only recovery ObjectStore. The chart pins a public multi-architecture image digest. The CronJob uses `concurrencyPolicy: Forbid` and a bounded job deadline. Its logs contain the JSON result; failed runs have nonzero exit status. **The chart does not yet provide durable report storage, missed-run alerts, or fleet policy.**
110
+
111
+ ## Safety boundary
112
+
113
+ - The generated recovery Cluster never copies `spec.plugins`, so it is not configured to archive WAL back to the source bucket. Set `recoveryObjectStore` to a separately named `ObjectStore` with read-only archive credentials. The tool compares its destination and endpoint with the source store; it cannot verify credential permissions, so test that writes are denied before relying on this boundary. Omitting this field uses the source `ObjectStore` for compatibility.
114
+ - A recovery timeout includes a machine-readable `failureReason` when the recovery Pod logs identify archive access denial or unavailable WAL. It does not copy raw logs into the report; inspect Pod and operator logs for detail.
115
+ - The drill uses the same namespace because the plugin's `ObjectStore` and its credentials are namespace-scoped. The source Cluster object is never modified.
116
+ - SQL checks run in a `BEGIN READ ONLY` transaction. Queries are limited to one `SELECT` or `WITH` statement without semicolons. Only provide trusted SQL; this is a guard against mistakes, not a security sandbox.
117
+ - By default, the tool deletes the drill Cluster even when recovery or a check fails. `retainOnFailure` leaves it for investigation and can incur storage costs. Cleanup failures turn the result red.
118
+ - The tool checks that PVCs labeled for the drill cluster disappear after Cluster deletion and fails the report if they remain. Kubernetes PersistentVolume cleanup depends on the storage class; inspect cloud volumes after a first drill.
119
+ - The planned SaaS boundary is metadata only. This repository contains no telemetry upload or hosted service.
120
+
121
+ ## Contributing and upstream distribution
122
+
123
+ Run `PYTHONPATH=src python3 -m unittest discover -s tests -v`. See [CONTRIBUTING.md](CONTRIBUTING.md) for how to add a recovery path, the live integration test gate, and the Krew package plan. We will submit only a tested, released package that stands on its own. Upstream acceptance is not assumed.
124
+
125
+ [UPSTREAM.md](UPSTREAM.md) records the specific repositories, proposed contribution value, and submission gates.
126
+
127
+ ## Recovery references
128
+
129
+ - [CloudNativePG backup guidance](https://cloudnative-pg.io/docs/devel/backup/)
130
+ - [CloudNativePG recovery documentation](https://cloudnative-pg.io/docs/devel/recovery/)
131
+ - [Barman Cloud recovery example](https://cloudnative-pg.io/plugin-barman-cloud/docs/next/concepts/)
@@ -0,0 +1,95 @@
1
+ # cnpg-drill
2
+
3
+ **Prove that a CloudNativePG backup can restore.** `cnpg-drill` creates a disposable CloudNativePG cluster from the source cluster's Barman Cloud object store, waits for PostgreSQL to become ready, runs read-only SQL assertions, emits a JSON result, and deletes the drill cluster. It can run locally or on a schedule as a Kubernetes CronJob.
4
+
5
+ The core is free and works without an account or external service. It does not make backups or copy database contents to a vendor. The proposed hosted service in [RESEARCH.md](RESEARCH.md) is separate and has not been built.
6
+
7
+ ## Status
8
+
9
+ **Early alpha.** Unit tests cover manifest safety and execution state. A disposable [local integration run](integration/local/README.md) on v0.1.0 proved full restore, PITR, read-only archive access, denied-WAL reporting, failed-check reporting, and Cluster/PVC cleanup on one version matrix. The v0.1.1 release archive and public chart 0.1.2 also restored a fresh backup and checked a row in an application database. A separate single-node pgEdge chart restore passed SQL checks in its `app` database. Validate against your operator, PostgreSQL image, storage, object store, and Barman plugin versions before treating a pass as disaster recovery assurance.
10
+
11
+ The [v0.1.2 release](https://github.com/danielgaskins/cnpg-drill/releases/tag/v0.1.2) carries literal AWS checksum compatibility settings into recovery clusters without copying source credentials. It includes the MIT-licensed kubectl plugin archive and checksum. The [default Krew index submission](https://github.com/kubernetes-sigs/krew-index/pull/6373) and [custom Krew index submission](https://github.com/ishantanu/awesome-kubectl-plugins/pull/44) currently package v0.1.0 and are under review. [kubetools](https://github.com/collabnix/kubetools/pull/431) is reviewing a Backup Tools listing.
12
+
13
+ The [Helm chart is listed on Artifact Hub](https://artifacthub.io/packages/helm/cnpg-drill/cnpg-drill) as a Verified Publisher package. Chart version `0.1.2` uses the v0.1.1 application image and includes the publisher's [website](https://danielgaskins.com/).
14
+
15
+ ## What it supports
16
+
17
+ - One source CloudNativePG cluster with one enabled Barman Cloud plugin and a named `ObjectStore` in the same namespace.
18
+ - An optional separate `recoveryObjectStore` with credentials limited to archive reads. Preflight checks that its destination and endpoint match the source store before creating a Cluster.
19
+ - Full restore from the latest completed Barman plugin `Backup` resource, or PITR to an explicit RFC3339 `targetTime`. The selected backup ID is pinned and included in the report; a stale backup fails preflight.
20
+ - Single-instance drill cluster using the source's PostgreSQL image, storage size/class, optional WAL storage, and resource requests.
21
+ - Copies literal AWS checksum compatibility settings from the source Cluster for S3-compatible archive restores; it does not copy writer credentials or arbitrary environment variables.
22
+ - Read-only SQL checks in `postgres` or a named application database. Exact expected values make reports easy to audit.
23
+ - Optional timestamp freshness assertions (`maxAgeSeconds`) for an application-defined recovery point. This measures age of that application's data, not the database WAL RPO.
24
+ - A JSON report on stdout and optionally in a local file. Query outputs are hashed rather than logged. Exit code 0 means all checks, Cluster deletion, and PVC cleanup passed; 1 means the drill failed; 2 means invalid input or a failed preflight.
25
+ - Scheduled execution with the Helm CronJob in `deploy/helm/cnpg-drill`.
26
+
27
+ The CLI refuses source clusters with tablespaces or recovery bootstrap because a naive clone could produce an incomplete or misleading pass. Volume snapshots, cross-namespace restores, cross-region recovery, custom Postgres extension image changes, and multi-cluster fleet management are future work.
28
+
29
+ ## Install and use
30
+
31
+ For a first run on your own cluster, follow the [first-run guide](docs/FIRST-RUN.md).
32
+
33
+ Requirements: Python 3.10+, `kubectl` in `PATH`, CloudNativePG and the Barman Cloud plugin installed in the target cluster, and Kubernetes access to read the source Cluster and ObjectStore, create/get/delete a drill Cluster, and exec into its PostgreSQL pod.
34
+
35
+ ```bash
36
+ python3 -m pip install -e .
37
+ cnpg-drill plan --config examples/drill.json
38
+ cnpg-drill run --config examples/drill.json --report reports/latest.json
39
+ ```
40
+
41
+ `plan` reads the source configuration and prints the exact Cluster manifest it would create. Review that manifest before running the first drill. The example's table-count assertion is illustrative; replace it with an invariant from your own application. A minimal check is `SELECT 1`, but it only proves connection to the recovered server, not useful application data.
42
+
43
+ Example `drill.json`:
44
+
45
+ ```json
46
+ {
47
+ "namespace": "production",
48
+ "cluster": "app-db",
49
+ "recoveryObjectStore": "app-db-recovery-readonly",
50
+ "timeoutSeconds": 1800,
51
+ "maxBackupAgeSeconds": 691200,
52
+ "checks": [
53
+ {"name": "postgres-ready", "query": "SELECT 1", "expected": "1"},
54
+ {"name": "orders-exist", "database": "app", "query": "SELECT count(*) > 0 FROM public.orders", "expected": "t"},
55
+ {"name": "recent-order", "database": "app", "query": "SELECT max(created_at)::timestamptz FROM public.orders", "maxAgeSeconds": 86400}
56
+ ]
57
+ }
58
+ ```
59
+
60
+ For PITR, add `"targetTime": "2026-09-29T10:00:00Z"`. Pick a time inside the backup and WAL retention window. A passing PITR check proves only that target; use a second drill for the latest recovery path.
61
+
62
+ ## Scheduled drill
63
+
64
+ The Helm chart creates a namespace-scoped ServiceAccount, Role, RoleBinding, ConfigMap, and CronJob. Set `cluster`, `checks`, and the other settings in a values file and install it **in the same namespace as the source Cluster and ObjectStore**. The chart defaults to suspension so installation does not immediately start a restore.
65
+
66
+ The example image build pins kubectl 1.36.4, suitable for Kubernetes 1.35–1.37 under the project's [version skew policy](https://kubernetes.io/releases/). Set `KUBECTL_VERSION` at image build time for another supported cluster version.
67
+
68
+ ```bash
69
+ helm install cnpg-drill oci://ghcr.io/danielgaskins/charts/cnpg-drill \
70
+ --version 0.1.2 -n production -f drill-values.yaml
71
+ ```
72
+
73
+ See the [first-run guide](docs/FIRST-RUN.md) and [chart values](deploy/helm/cnpg-drill/README.md) for a suspended first run with a separate read-only recovery ObjectStore. The chart pins a public multi-architecture image digest. The CronJob uses `concurrencyPolicy: Forbid` and a bounded job deadline. Its logs contain the JSON result; failed runs have nonzero exit status. **The chart does not yet provide durable report storage, missed-run alerts, or fleet policy.**
74
+
75
+ ## Safety boundary
76
+
77
+ - The generated recovery Cluster never copies `spec.plugins`, so it is not configured to archive WAL back to the source bucket. Set `recoveryObjectStore` to a separately named `ObjectStore` with read-only archive credentials. The tool compares its destination and endpoint with the source store; it cannot verify credential permissions, so test that writes are denied before relying on this boundary. Omitting this field uses the source `ObjectStore` for compatibility.
78
+ - A recovery timeout includes a machine-readable `failureReason` when the recovery Pod logs identify archive access denial or unavailable WAL. It does not copy raw logs into the report; inspect Pod and operator logs for detail.
79
+ - The drill uses the same namespace because the plugin's `ObjectStore` and its credentials are namespace-scoped. The source Cluster object is never modified.
80
+ - SQL checks run in a `BEGIN READ ONLY` transaction. Queries are limited to one `SELECT` or `WITH` statement without semicolons. Only provide trusted SQL; this is a guard against mistakes, not a security sandbox.
81
+ - By default, the tool deletes the drill Cluster even when recovery or a check fails. `retainOnFailure` leaves it for investigation and can incur storage costs. Cleanup failures turn the result red.
82
+ - The tool checks that PVCs labeled for the drill cluster disappear after Cluster deletion and fails the report if they remain. Kubernetes PersistentVolume cleanup depends on the storage class; inspect cloud volumes after a first drill.
83
+ - The planned SaaS boundary is metadata only. This repository contains no telemetry upload or hosted service.
84
+
85
+ ## Contributing and upstream distribution
86
+
87
+ Run `PYTHONPATH=src python3 -m unittest discover -s tests -v`. See [CONTRIBUTING.md](CONTRIBUTING.md) for how to add a recovery path, the live integration test gate, and the Krew package plan. We will submit only a tested, released package that stands on its own. Upstream acceptance is not assumed.
88
+
89
+ [UPSTREAM.md](UPSTREAM.md) records the specific repositories, proposed contribution value, and submission gates.
90
+
91
+ ## Recovery references
92
+
93
+ - [CloudNativePG backup guidance](https://cloudnative-pg.io/docs/devel/backup/)
94
+ - [CloudNativePG recovery documentation](https://cloudnative-pg.io/docs/devel/recovery/)
95
+ - [Barman Cloud recovery example](https://cloudnative-pg.io/plugin-barman-cloud/docs/next/concepts/)
@@ -0,0 +1,24 @@
1
+ [build-system]
2
+ requires = ["setuptools>=68"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "cnpg-drill"
7
+ version = "0.1.2"
8
+ description = "Prove CloudNativePG Barman backups restore into an isolated cluster"
9
+ readme = "README.md"
10
+ requires-python = ">=3.10"
11
+ license = {file = "LICENSE"}
12
+ authors = [{name = "cnpg-drill contributors"}]
13
+ keywords = ["cloudnativepg", "postgresql", "kubernetes", "backup", "restore"]
14
+ classifiers = [
15
+ "Programming Language :: Python :: 3",
16
+ "License :: OSI Approved :: MIT License",
17
+ "Operating System :: OS Independent",
18
+ ]
19
+
20
+ [project.scripts]
21
+ cnpg-drill = "cnpg_drill.cli:main"
22
+
23
+ [tool.setuptools.packages.find]
24
+ where = ["src"]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,3 @@
1
+ """CloudNativePG recovery drill."""
2
+
3
+ __version__ = "0.1.2"
@@ -0,0 +1,43 @@
1
+ """Command-line interface for local use and Kubernetes CronJobs."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import argparse
6
+ import json
7
+ import sys
8
+ from pathlib import Path
9
+
10
+ from . import __version__
11
+ from .core import Config, DrillError, Kubectl, prepare, run_drill
12
+
13
+
14
+ def main(argv: list[str] | None = None) -> int:
15
+ parser = argparse.ArgumentParser(prog="cnpg-drill", description="Verify a CloudNativePG backup with a disposable recovery cluster")
16
+ parser.add_argument("--version", action="version", version=__version__)
17
+ parser.add_argument("--kubectl", default="kubectl", help="kubectl executable path")
18
+ sub = parser.add_subparsers(dest="command", required=True)
19
+ for command in ("plan", "run"):
20
+ item = sub.add_parser(command)
21
+ item.add_argument("--config", required=True, type=Path, help="JSON drill config")
22
+ if command == "run":
23
+ item.add_argument("--report", type=Path, help="write JSON report to this path")
24
+ args = parser.parse_args(argv)
25
+ try:
26
+ config = Config.from_dict(json.loads(args.config.read_text(encoding="utf-8")))
27
+ client = Kubectl(args.kubectl)
28
+ if args.command == "plan":
29
+ print(json.dumps(prepare(client, config), indent=2, sort_keys=True))
30
+ return 0
31
+ report = run_drill(client, config)
32
+ encoded = json.dumps(report, indent=2, sort_keys=True)
33
+ print(encoded)
34
+ if args.report:
35
+ args.report.write_text(encoded + "\n", encoding="utf-8")
36
+ return 0 if report["status"] == "passed" else (2 if report.get("phase") == "preflight" else 1)
37
+ except (DrillError, OSError, json.JSONDecodeError) as exc:
38
+ print(f"cnpg-drill: {exc}", file=sys.stderr)
39
+ return 2
40
+
41
+
42
+ if __name__ == "__main__":
43
+ raise SystemExit(main())
@@ -0,0 +1,393 @@
1
+ """Recovery planning and execution, with a deliberately small Kubernetes surface."""
2
+
3
+ from __future__ import annotations
4
+
5
+ import copy
6
+ import datetime as dt
7
+ import hashlib
8
+ import json
9
+ import re
10
+ import secrets
11
+ import subprocess
12
+ import time
13
+ from dataclasses import dataclass
14
+ from typing import Any, Callable
15
+
16
+
17
+ PLUGIN = "barman-cloud.cloudnative-pg.io"
18
+ API = "clusters.postgresql.cnpg.io"
19
+ DNS_NAME = re.compile(r"^[a-z0-9]([-a-z0-9]*[a-z0-9])?$")
20
+
21
+
22
+ class DrillError(RuntimeError):
23
+ """A recoverable validation or drill failure."""
24
+
25
+
26
+ @dataclass(frozen=True)
27
+ class Check:
28
+ name: str
29
+ query: str
30
+ expected: str | None = None
31
+ max_age_seconds: int | None = None
32
+ database: str = "postgres"
33
+
34
+
35
+ @dataclass(frozen=True)
36
+ class Config:
37
+ namespace: str
38
+ cluster: str
39
+ recovery_object_store: str | None = None
40
+ timeout_seconds: int = 1800
41
+ poll_seconds: int = 5
42
+ max_backup_age_seconds: int = 691200
43
+ target_time: str | None = None
44
+ checks: tuple[Check, ...] = (Check("connection", "SELECT 1", "1"),)
45
+ retain_on_failure: bool = False
46
+
47
+ @classmethod
48
+ def from_dict(cls, raw: dict[str, Any]) -> Config:
49
+ if not isinstance(raw, dict):
50
+ raise DrillError("Config must be a JSON object")
51
+ allowed = {"namespace", "cluster", "recoveryObjectStore", "timeoutSeconds", "pollSeconds", "maxBackupAgeSeconds", "targetTime", "checks", "retainOnFailure"}
52
+ unknown = set(raw) - allowed
53
+ if unknown:
54
+ raise DrillError(f"Unknown config keys: {', '.join(sorted(unknown))}")
55
+ namespace, cluster = raw.get("namespace"), raw.get("cluster")
56
+ for label, value in (("namespace", namespace), ("cluster", cluster)):
57
+ if not isinstance(value, str) or not DNS_NAME.fullmatch(value) or len(value) > 63:
58
+ raise DrillError(f"{label} must be a Kubernetes DNS label")
59
+ recovery_object_store = raw.get("recoveryObjectStore")
60
+ if recovery_object_store is not None and (not isinstance(recovery_object_store, str) or not DNS_NAME.fullmatch(recovery_object_store) or len(recovery_object_store) > 63):
61
+ raise DrillError("recoveryObjectStore must be a Kubernetes DNS label")
62
+ timeout = raw.get("timeoutSeconds", 1800)
63
+ poll = raw.get("pollSeconds", 5)
64
+ max_backup_age = raw.get("maxBackupAgeSeconds", 691200)
65
+ if type(timeout) is not int or not 30 <= timeout <= 86400:
66
+ raise DrillError("timeoutSeconds must be 30..86400")
67
+ if type(poll) is not int or not 1 <= poll <= 60:
68
+ raise DrillError("pollSeconds must be 1..60")
69
+ if type(max_backup_age) is not int or not 1 <= max_backup_age <= 31536000:
70
+ raise DrillError("maxBackupAgeSeconds must be 1..31536000")
71
+ target = raw.get("targetTime")
72
+ if target is not None:
73
+ if not isinstance(target, str):
74
+ raise DrillError("targetTime must be an RFC3339 timestamp")
75
+ try:
76
+ parsed = dt.datetime.fromisoformat(target.replace("Z", "+00:00"))
77
+ except ValueError as exc:
78
+ raise DrillError("targetTime must be an RFC3339 timestamp") from exc
79
+ if parsed.tzinfo is None:
80
+ raise DrillError("targetTime must include a timezone")
81
+ if parsed > dt.datetime.now(dt.timezone.utc):
82
+ raise DrillError("targetTime cannot be in the future")
83
+ checks_raw = raw.get("checks", [{"name": "connection", "query": "SELECT 1", "expected": "1"}])
84
+ if not isinstance(checks_raw, list) or not 1 <= len(checks_raw) <= 50:
85
+ raise DrillError("checks must be a list of 1..50 checks")
86
+ checks: list[Check] = []
87
+ for item in checks_raw:
88
+ if not isinstance(item, dict) or not {"name", "query"} <= set(item) or set(item) - {"name", "query", "expected", "maxAgeSeconds", "database"}:
89
+ raise DrillError("Each check needs name, query, and one comparison")
90
+ if ("expected" in item) == ("maxAgeSeconds" in item):
91
+ raise DrillError("Each check needs exactly one of expected or maxAgeSeconds")
92
+ name, query, expected = item["name"], item["query"], item.get("expected")
93
+ max_age = item.get("maxAgeSeconds")
94
+ database = item.get("database", "postgres")
95
+ if not isinstance(name, str) or not name or len(name) > 80:
96
+ raise DrillError("Check name must be 1..80 characters")
97
+ if "expected" in item and (not isinstance(expected, str) or len(expected) > 4096):
98
+ raise DrillError("Check expected must be a string up to 4096 characters")
99
+ if "maxAgeSeconds" in item and (type(max_age) is not int or not 1 <= max_age <= 31536000):
100
+ raise DrillError("maxAgeSeconds must be 1..31536000")
101
+ if not isinstance(query, str) or len(query) > 8192 or not re.match(r"^\s*(SELECT|WITH)\b", query, re.I) or ";" in query:
102
+ raise DrillError(f"Check {name}: query must be one SELECT/WITH statement without semicolons")
103
+ if not isinstance(database, str) or not re.fullmatch(r"[A-Za-z_][A-Za-z0-9_-]{0,62}", database):
104
+ raise DrillError(f"Check {name}: database must be a simple PostgreSQL database name")
105
+ checks.append(Check(name, query.strip(), expected, max_age, database))
106
+ retained = raw.get("retainOnFailure", False)
107
+ if type(retained) is not bool:
108
+ raise DrillError("retainOnFailure must be boolean")
109
+ return cls(namespace, cluster, recovery_object_store, timeout, poll, max_backup_age, target, tuple(checks), retained)
110
+
111
+
112
+ class Kubectl:
113
+ def __init__(self, executable: str = "kubectl", invoke: Callable[..., subprocess.CompletedProcess[str]] = subprocess.run):
114
+ self.executable = executable
115
+ self.invoke = invoke
116
+
117
+ def call(self, args: list[str], *, input_text: str | None = None, timeout: int = 60) -> str:
118
+ try:
119
+ result = self.invoke([self.executable, *args], input=input_text, text=True, capture_output=True, timeout=timeout, check=False)
120
+ except (FileNotFoundError, subprocess.TimeoutExpired) as exc:
121
+ raise DrillError(f"kubectl unavailable or timed out: {exc}") from exc
122
+ if result.returncode:
123
+ raise DrillError(f"kubectl {' '.join(args[:4])} failed: {result.stderr.strip()[:1000]}")
124
+ return result.stdout
125
+
126
+ def get(self, resource: str, name: str, namespace: str) -> dict[str, Any]:
127
+ return json.loads(self.call(["-n", namespace, "get", resource, name, "-o", "json"]))
128
+
129
+ def list_pvcs(self, cluster: str, namespace: str) -> list[str]:
130
+ data = json.loads(self.call(["-n", namespace, "get", "pvc", "-l", f"cnpg.io/cluster={cluster}", "-o", "json"]))
131
+ return [item["metadata"]["name"] for item in data.get("items", [])]
132
+
133
+ def recovery_logs(self, cluster: str, namespace: str) -> list[str]:
134
+ data = json.loads(self.call(["-n", namespace, "get", "pods", "-l", f"cnpg.io/cluster={cluster}", "-o", "json"]))
135
+ logs = []
136
+ pods = sorted(data.get("items", []), key=lambda p: p.get("metadata", {}).get("creationTimestamp", ""), reverse=True)
137
+ for pod in pods[:5]:
138
+ name = pod.get("metadata", {}).get("name", "")
139
+ if "full-recovery" not in name:
140
+ continue
141
+ for container in ("plugin-barman-cloud", "full-recovery"):
142
+ try:
143
+ logs.append(self.call(["-n", namespace, "logs", name, "-c", container, "--tail=200"], timeout=20))
144
+ except DrillError:
145
+ continue
146
+ return logs
147
+
148
+ def list_backups(self, namespace: str) -> list[dict[str, Any]]:
149
+ data = json.loads(self.call(["-n", namespace, "get", "backups.postgresql.cnpg.io", "-o", "json"]))
150
+ return data.get("items", [])
151
+
152
+ def create(self, manifest: dict[str, Any], namespace: str) -> None:
153
+ self.call(["-n", namespace, "create", "-f", "-"], input_text=json.dumps(manifest))
154
+
155
+ def delete(self, resource: str, name: str, namespace: str, *, timeout: int = 120) -> None:
156
+ self.call(["-n", namespace, "delete", resource, name, "--ignore-not-found=true", "--wait=true", f"--timeout={timeout}s"], timeout=timeout + 15)
157
+
158
+ def exec_query(self, pod: str, namespace: str, query: str, database: str = "postgres") -> str:
159
+ # One read-only transaction. The config parser rejects semicolons in query.
160
+ sql = f"BEGIN READ ONLY; {query}; COMMIT;"
161
+ return self.call(["-n", namespace, "exec", pod, "-c", "postgres", "--", "psql", "-X", "-qAt", "-v", "ON_ERROR_STOP=1", "-U", "postgres", "-d", database, "-c", sql], timeout=60).strip()
162
+
163
+
164
+ def choose_backup(backups: list[dict[str, Any]], config: Config, object_name: str | None = None) -> tuple[dict[str, Any], int]:
165
+ now = dt.datetime.now(dt.timezone.utc)
166
+ target = dt.datetime.fromisoformat(config.target_time.replace("Z", "+00:00")) if config.target_time else now
167
+ candidates: list[tuple[dt.datetime, dict[str, Any]]] = []
168
+ for backup in backups:
169
+ if backup.get("spec", {}).get("cluster", {}).get("name") != config.cluster:
170
+ continue
171
+ status = backup.get("status", {})
172
+ if status.get("phase") != "completed" or not status.get("backupId"):
173
+ continue
174
+ if (status.get("method") or backup.get("spec", {}).get("method")) != "plugin":
175
+ continue
176
+ if backup.get("spec", {}).get("pluginConfiguration", {}).get("name") not in (None, PLUGIN):
177
+ continue
178
+ backup_object = backup.get("spec", {}).get("pluginConfiguration", {}).get("parameters", {}).get("barmanObjectName")
179
+ if object_name and backup_object and backup_object != object_name:
180
+ continue
181
+ try:
182
+ stopped = dt.datetime.fromisoformat(status["stoppedAt"].replace("Z", "+00:00"))
183
+ except (KeyError, TypeError, ValueError):
184
+ continue
185
+ if stopped.tzinfo and stopped <= target:
186
+ candidates.append((stopped, backup))
187
+ if not candidates:
188
+ raise DrillError("No completed Barman plugin Backup with an ID exists before the recovery target")
189
+ stopped, chosen = max(candidates, key=lambda item: item[0])
190
+ age = int((now - stopped).total_seconds())
191
+ if age > config.max_backup_age_seconds:
192
+ raise DrillError(f"Latest eligible Backup is {age}s old, above maxBackupAgeSeconds={config.max_backup_age_seconds}")
193
+ return chosen, age
194
+
195
+
196
+ def build_manifest(source: dict[str, Any], config: Config, *, name: str, backup: dict[str, Any] | None = None) -> dict[str, Any]:
197
+ spec = source.get("spec", {})
198
+ if spec.get("tablespaces"):
199
+ raise DrillError("Tablespace clusters are not supported in v0.1; refusing incomplete restore")
200
+ bootstrap = spec.get("bootstrap") or {}
201
+ if bootstrap.get("recovery") or bootstrap.get("pg_basebackup"):
202
+ raise DrillError("Replica or recovery-source clusters are not supported in v0.1")
203
+ if bootstrap and "initdb" not in bootstrap:
204
+ raise DrillError("Only default or initdb bootstrap sources are supported in v0.1")
205
+ plugins = [p for p in spec.get("plugins", []) if p.get("name") == PLUGIN and p.get("enabled", True)]
206
+ if len(plugins) != 1:
207
+ raise DrillError("Source needs one enabled Barman Cloud plugin")
208
+ params = plugins[0].get("parameters", {})
209
+ object_name = params.get("barmanObjectName")
210
+ if not isinstance(object_name, str) or not DNS_NAME.fullmatch(object_name):
211
+ raise DrillError("Barman plugin needs a valid barmanObjectName")
212
+ storage = spec.get("storage")
213
+ if not isinstance(storage, dict) or not storage.get("size"):
214
+ raise DrillError("Source Cluster needs storage.size")
215
+ if len(name) > 63 or not DNS_NAME.fullmatch(name) or name == config.cluster:
216
+ raise DrillError("Invalid or unsafe drill cluster name")
217
+ recovery: dict[str, Any] = {"source": "backup-source"}
218
+ if backup:
219
+ recovery["recoveryTarget"] = {"backupID": backup["status"]["backupId"]}
220
+ initdb = bootstrap.get("initdb") or {}
221
+ for key in ("database", "owner", "secret"):
222
+ if key in initdb:
223
+ recovery[key] = copy.deepcopy(initdb[key])
224
+ if config.target_time:
225
+ recovery.setdefault("recoveryTarget", {})["targetTime"] = config.target_time
226
+ external = {"name": "backup-source", "plugin": {"name": PLUGIN, "enabled": True, "parameters": {"barmanObjectName": config.recovery_object_store or object_name, "serverName": params.get("serverName") or config.cluster}}}
227
+ output: dict[str, Any] = {
228
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": "Cluster",
229
+ "metadata": {"name": name, "namespace": config.namespace, "labels": {"app.kubernetes.io/managed-by": "cnpg-drill", "cnpg-drill.dev/source": config.cluster}, "annotations": {"cnpg-drill.dev/source-uid": source.get("metadata", {}).get("uid", ""), "cnpg-drill.dev/backup-name": backup.get("metadata", {}).get("name", "") if backup else ""}},
230
+ "spec": {"instances": 1, "storage": copy.deepcopy(storage), "bootstrap": {"recovery": recovery}, "externalClusters": [external]},
231
+ }
232
+ for key in ("imageName", "imageCatalogRef", "walStorage", "resources"):
233
+ if key in spec:
234
+ output["spec"][key] = copy.deepcopy(spec[key])
235
+ # S3-compatible stores may need these boto3 settings during recovery too.
236
+ # Never copy arbitrary source env: it can contain writer credentials.
237
+ recovery_env = [
238
+ {"name": item["name"], "value": item["value"]} for item in spec.get("env", [])
239
+ if isinstance(item, dict)
240
+ and item.get("name") in {"AWS_REQUEST_CHECKSUM_CALCULATION", "AWS_RESPONSE_CHECKSUM_VALIDATION", "AWS_NO_CHUNKED_ENCODING"}
241
+ and isinstance(item.get("value"), str)
242
+ and "valueFrom" not in item
243
+ ]
244
+ if recovery_env:
245
+ output["spec"]["env"] = recovery_env
246
+ # Deliberately do not copy spec.plugins: the recovered cluster must not archive to source.
247
+ return output
248
+
249
+
250
+ def new_name(source_name: str) -> str:
251
+ suffix = dt.datetime.now(dt.timezone.utc).strftime("%m%d%H%M") + "-" + secrets.token_hex(3)
252
+ return f"drill-{source_name[:48-len(suffix)]}-{suffix}"
253
+
254
+
255
+ def recovery_failure_reason(logs: list[str]) -> str | None:
256
+ """Classify recovery logs without exposing archive paths, credentials, or SQL."""
257
+ messages = []
258
+ for log in logs:
259
+ for line in log.splitlines()[-200:]:
260
+ try:
261
+ entry = json.loads(line)
262
+ except json.JSONDecodeError:
263
+ continue
264
+ if not isinstance(entry, dict):
265
+ continue
266
+ record = entry.get("record")
267
+ if isinstance(record, dict):
268
+ messages.append(str(record.get("message", "")))
269
+ messages.append(str(entry.get("error", "")))
270
+ messages.append(str(entry.get("msg", "")))
271
+ combined = "\n".join(messages).lower()
272
+ if re.search(r"access.?denied|permission.?denied|forbidden|\b403\b", combined):
273
+ return "archive_access_denied"
274
+ if re.search(r"(wal|archive).{0,100}(not found|no such key|missing|unavailable|does not exist)|requested wal segment.{0,100}removed", combined):
275
+ return "wal_unavailable"
276
+ if re.search(r"restore error|error while restoring a backup|fatal", combined):
277
+ return "recovery_process_failed"
278
+ return None
279
+
280
+
281
+ def prepare(client: Kubectl, config: Config, name: str | None = None) -> dict[str, Any]:
282
+ source = client.get(API, config.cluster, config.namespace)
283
+ if source.get("metadata", {}).get("deletionTimestamp"):
284
+ raise DrillError("Source Cluster is being deleted")
285
+ source_plugins = [p for p in source.get("spec", {}).get("plugins", []) if p.get("name") == PLUGIN and p.get("enabled", True)]
286
+ object_name = source_plugins[0].get("parameters", {}).get("barmanObjectName") if len(source_plugins) == 1 else None
287
+ backup, age = choose_backup(client.list_backups(config.namespace), config, object_name)
288
+ manifest = build_manifest(source, config, name=name or new_name(config.cluster), backup=backup)
289
+ manifest["metadata"]["annotations"]["cnpg-drill.dev/backup-age-seconds"] = str(age)
290
+ if not object_name:
291
+ raise DrillError("Source needs one enabled Barman Cloud plugin")
292
+ source_store = client.get("objectstores.barmancloud.cnpg.io", object_name, config.namespace)
293
+ recovery_name = config.recovery_object_store or object_name
294
+ if recovery_name != object_name:
295
+ recovery_store = client.get("objectstores.barmancloud.cnpg.io", recovery_name, config.namespace)
296
+ source_config = source_store.get("spec", {}).get("configuration", {})
297
+ recovery_config = recovery_store.get("spec", {}).get("configuration", {})
298
+ for field in ("destinationPath", "endpointURL"):
299
+ if not source_config.get(field) or source_config.get(field) != recovery_config.get(field):
300
+ raise DrillError(f"Recovery ObjectStore {field} must match source ObjectStore")
301
+ return manifest
302
+
303
+
304
+ def run_drill(client: Kubectl, config: Config, *, sleep: Callable[[float], None] = time.sleep, monotonic: Callable[[], float] = time.monotonic) -> dict[str, Any]:
305
+ started_at = dt.datetime.now(dt.timezone.utc)
306
+ started = monotonic()
307
+ result: dict[str, Any] = {"schemaVersion": 1, "source": f"{config.namespace}/{config.cluster}", "startedAt": started_at.isoformat(), "targetTime": config.target_time, "checks": [], "status": "error"}
308
+ name: str | None = None
309
+ created = False
310
+ try:
311
+ manifest = prepare(client, config)
312
+ name = manifest["metadata"]["name"]
313
+ result["drillCluster"] = name
314
+ result["sourceClusterUID"] = manifest["metadata"]["annotations"]["cnpg-drill.dev/source-uid"]
315
+ result["backupName"] = manifest["metadata"]["annotations"]["cnpg-drill.dev/backup-name"]
316
+ result["backupID"] = manifest["spec"]["bootstrap"]["recovery"]["recoveryTarget"]["backupID"]
317
+ result["backupAgeSeconds"] = int(manifest["metadata"]["annotations"]["cnpg-drill.dev/backup-age-seconds"])
318
+ created = True
319
+ # A create request can reach the API even if the client loses its response.
320
+ # In that case, still try to remove the uniquely named drill Cluster.
321
+ client.create(manifest, config.namespace)
322
+ recovery_started = monotonic()
323
+ deadline = started + config.timeout_seconds
324
+ primary = None
325
+ while monotonic() < deadline:
326
+ cluster = client.get(API, name, config.namespace)
327
+ status = cluster.get("status", {})
328
+ if status.get("readyInstances", 0) >= 1 and status.get("currentPrimary"):
329
+ primary = status["currentPrimary"]
330
+ break
331
+ sleep(config.poll_seconds)
332
+ if not primary:
333
+ try:
334
+ reason = recovery_failure_reason(client.recovery_logs(name, config.namespace))
335
+ except DrillError:
336
+ reason = None
337
+ result["failureReason"] = reason or "recovery_timeout"
338
+ hints = {
339
+ "archive_access_denied": "Check recovery ObjectStore permissions to read both base backup and WAL objects.",
340
+ "wal_unavailable": "Check that the required WAL segment exists and is readable in the recovery archive.",
341
+ "recovery_process_failed": "Inspect the drill recovery Pod and operator logs for the underlying restore error.",
342
+ "recovery_timeout": "Inspect the drill recovery Pod and operator logs for the cause.",
343
+ }
344
+ raise DrillError(f"Recovery did not become ready within {config.timeout_seconds}s. {hints[result['failureReason']]}")
345
+ result["recoverySeconds"] = round(monotonic() - recovery_started, 2)
346
+ for check in config.checks:
347
+ observed = client.exec_query(primary, config.namespace, check.query, check.database)
348
+ check_result: dict[str, Any] = {"name": check.name, "database": check.database, "observedSha256": hashlib.sha256(observed.encode()).hexdigest()}
349
+ if check.max_age_seconds is not None:
350
+ try:
351
+ observed_time = dt.datetime.fromisoformat(observed.replace("Z", "+00:00"))
352
+ if observed_time.tzinfo is None:
353
+ raise ValueError("timestamp lacks timezone")
354
+ age = (dt.datetime.now(dt.timezone.utc) - observed_time).total_seconds()
355
+ check_result["ageSeconds"] = round(age, 2)
356
+ check_result["maxAgeSeconds"] = check.max_age_seconds
357
+ passed = 0 <= age <= check.max_age_seconds
358
+ except ValueError:
359
+ passed = False
360
+ check_result["error"] = "query did not return a timestamp with timezone"
361
+ else:
362
+ passed = observed == check.expected
363
+ check_result["passed"] = passed
364
+ result["checks"].append(check_result)
365
+ if not passed:
366
+ raise DrillError(f"Check {check.name!r} did not meet its assertion")
367
+ result["status"] = "passed"
368
+ except (DrillError, ValueError, KeyError) as exc:
369
+ result["error"] = str(exc)
370
+ result["status"] = "failed"
371
+ if not created:
372
+ result["phase"] = "preflight"
373
+ finally:
374
+ if created and name and not (result["status"] == "failed" and config.retain_on_failure):
375
+ try:
376
+ client.delete(API, name, config.namespace)
377
+ cleanup_deadline = monotonic() + 60
378
+ remaining = client.list_pvcs(name, config.namespace)
379
+ while remaining and monotonic() < cleanup_deadline:
380
+ sleep(2)
381
+ remaining = client.list_pvcs(name, config.namespace)
382
+ if remaining:
383
+ raise DrillError(f"Drill PVCs still exist after Cluster deletion: {', '.join(remaining)}")
384
+ result["cleanup"] = "cluster-and-pvcs-deleted"
385
+ except DrillError as exc:
386
+ result["cleanup"] = "failed"
387
+ result["cleanupError"] = str(exc)
388
+ result["status"] = "failed"
389
+ elif created:
390
+ result["cleanup"] = "retained-for-investigation"
391
+ result["finishedAt"] = dt.datetime.now(dt.timezone.utc).isoformat()
392
+ result["durationSeconds"] = round(monotonic() - started, 2)
393
+ return result
@@ -0,0 +1,131 @@
1
+ Metadata-Version: 2.4
2
+ Name: cnpg-drill
3
+ Version: 0.1.2
4
+ Summary: Prove CloudNativePG Barman backups restore into an isolated cluster
5
+ Author: cnpg-drill contributors
6
+ License: MIT License
7
+
8
+ Copyright (c) 2026 cnpg-drill contributors
9
+
10
+ Permission is hereby granted, free of charge, to any person obtaining a copy
11
+ of this software and associated documentation files (the "Software"), to deal
12
+ in the Software without restriction, including without limitation the rights
13
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
14
+ copies of the Software, and to permit persons to whom the Software is
15
+ furnished to do so, subject to the following conditions:
16
+
17
+ The above copyright notice and this permission notice shall be included in all
18
+ copies or substantial portions of the Software.
19
+
20
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
21
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
22
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
23
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
24
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
25
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
26
+ SOFTWARE.
27
+
28
+ Keywords: cloudnativepg,postgresql,kubernetes,backup,restore
29
+ Classifier: Programming Language :: Python :: 3
30
+ Classifier: License :: OSI Approved :: MIT License
31
+ Classifier: Operating System :: OS Independent
32
+ Requires-Python: >=3.10
33
+ Description-Content-Type: text/markdown
34
+ License-File: LICENSE
35
+ Dynamic: license-file
36
+
37
+ # cnpg-drill
38
+
39
+ **Prove that a CloudNativePG backup can restore.** `cnpg-drill` creates a disposable CloudNativePG cluster from the source cluster's Barman Cloud object store, waits for PostgreSQL to become ready, runs read-only SQL assertions, emits a JSON result, and deletes the drill cluster. It can run locally or on a schedule as a Kubernetes CronJob.
40
+
41
+ The core is free and works without an account or external service. It does not make backups or copy database contents to a vendor. The proposed hosted service in [RESEARCH.md](RESEARCH.md) is separate and has not been built.
42
+
43
+ ## Status
44
+
45
+ **Early alpha.** Unit tests cover manifest safety and execution state. A disposable [local integration run](integration/local/README.md) on v0.1.0 proved full restore, PITR, read-only archive access, denied-WAL reporting, failed-check reporting, and Cluster/PVC cleanup on one version matrix. The v0.1.1 release archive and public chart 0.1.2 also restored a fresh backup and checked a row in an application database. A separate single-node pgEdge chart restore passed SQL checks in its `app` database. Validate against your operator, PostgreSQL image, storage, object store, and Barman plugin versions before treating a pass as disaster recovery assurance.
46
+
47
+ The [v0.1.2 release](https://github.com/danielgaskins/cnpg-drill/releases/tag/v0.1.2) carries literal AWS checksum compatibility settings into recovery clusters without copying source credentials. It includes the MIT-licensed kubectl plugin archive and checksum. The [default Krew index submission](https://github.com/kubernetes-sigs/krew-index/pull/6373) and [custom Krew index submission](https://github.com/ishantanu/awesome-kubectl-plugins/pull/44) currently package v0.1.0 and are under review. [kubetools](https://github.com/collabnix/kubetools/pull/431) is reviewing a Backup Tools listing.
48
+
49
+ The [Helm chart is listed on Artifact Hub](https://artifacthub.io/packages/helm/cnpg-drill/cnpg-drill) as a Verified Publisher package. Chart version `0.1.2` uses the v0.1.1 application image and includes the publisher's [website](https://danielgaskins.com/).
50
+
51
+ ## What it supports
52
+
53
+ - One source CloudNativePG cluster with one enabled Barman Cloud plugin and a named `ObjectStore` in the same namespace.
54
+ - An optional separate `recoveryObjectStore` with credentials limited to archive reads. Preflight checks that its destination and endpoint match the source store before creating a Cluster.
55
+ - Full restore from the latest completed Barman plugin `Backup` resource, or PITR to an explicit RFC3339 `targetTime`. The selected backup ID is pinned and included in the report; a stale backup fails preflight.
56
+ - Single-instance drill cluster using the source's PostgreSQL image, storage size/class, optional WAL storage, and resource requests.
57
+ - Copies literal AWS checksum compatibility settings from the source Cluster for S3-compatible archive restores; it does not copy writer credentials or arbitrary environment variables.
58
+ - Read-only SQL checks in `postgres` or a named application database. Exact expected values make reports easy to audit.
59
+ - Optional timestamp freshness assertions (`maxAgeSeconds`) for an application-defined recovery point. This measures age of that application's data, not the database WAL RPO.
60
+ - A JSON report on stdout and optionally in a local file. Query outputs are hashed rather than logged. Exit code 0 means all checks, Cluster deletion, and PVC cleanup passed; 1 means the drill failed; 2 means invalid input or a failed preflight.
61
+ - Scheduled execution with the Helm CronJob in `deploy/helm/cnpg-drill`.
62
+
63
+ The CLI refuses source clusters with tablespaces or recovery bootstrap because a naive clone could produce an incomplete or misleading pass. Volume snapshots, cross-namespace restores, cross-region recovery, custom Postgres extension image changes, and multi-cluster fleet management are future work.
64
+
65
+ ## Install and use
66
+
67
+ For a first run on your own cluster, follow the [first-run guide](docs/FIRST-RUN.md).
68
+
69
+ Requirements: Python 3.10+, `kubectl` in `PATH`, CloudNativePG and the Barman Cloud plugin installed in the target cluster, and Kubernetes access to read the source Cluster and ObjectStore, create/get/delete a drill Cluster, and exec into its PostgreSQL pod.
70
+
71
+ ```bash
72
+ python3 -m pip install -e .
73
+ cnpg-drill plan --config examples/drill.json
74
+ cnpg-drill run --config examples/drill.json --report reports/latest.json
75
+ ```
76
+
77
+ `plan` reads the source configuration and prints the exact Cluster manifest it would create. Review that manifest before running the first drill. The example's table-count assertion is illustrative; replace it with an invariant from your own application. A minimal check is `SELECT 1`, but it only proves connection to the recovered server, not useful application data.
78
+
79
+ Example `drill.json`:
80
+
81
+ ```json
82
+ {
83
+ "namespace": "production",
84
+ "cluster": "app-db",
85
+ "recoveryObjectStore": "app-db-recovery-readonly",
86
+ "timeoutSeconds": 1800,
87
+ "maxBackupAgeSeconds": 691200,
88
+ "checks": [
89
+ {"name": "postgres-ready", "query": "SELECT 1", "expected": "1"},
90
+ {"name": "orders-exist", "database": "app", "query": "SELECT count(*) > 0 FROM public.orders", "expected": "t"},
91
+ {"name": "recent-order", "database": "app", "query": "SELECT max(created_at)::timestamptz FROM public.orders", "maxAgeSeconds": 86400}
92
+ ]
93
+ }
94
+ ```
95
+
96
+ For PITR, add `"targetTime": "2026-09-29T10:00:00Z"`. Pick a time inside the backup and WAL retention window. A passing PITR check proves only that target; use a second drill for the latest recovery path.
97
+
98
+ ## Scheduled drill
99
+
100
+ The Helm chart creates a namespace-scoped ServiceAccount, Role, RoleBinding, ConfigMap, and CronJob. Set `cluster`, `checks`, and the other settings in a values file and install it **in the same namespace as the source Cluster and ObjectStore**. The chart defaults to suspension so installation does not immediately start a restore.
101
+
102
+ The example image build pins kubectl 1.36.4, suitable for Kubernetes 1.35–1.37 under the project's [version skew policy](https://kubernetes.io/releases/). Set `KUBECTL_VERSION` at image build time for another supported cluster version.
103
+
104
+ ```bash
105
+ helm install cnpg-drill oci://ghcr.io/danielgaskins/charts/cnpg-drill \
106
+ --version 0.1.2 -n production -f drill-values.yaml
107
+ ```
108
+
109
+ See the [first-run guide](docs/FIRST-RUN.md) and [chart values](deploy/helm/cnpg-drill/README.md) for a suspended first run with a separate read-only recovery ObjectStore. The chart pins a public multi-architecture image digest. The CronJob uses `concurrencyPolicy: Forbid` and a bounded job deadline. Its logs contain the JSON result; failed runs have nonzero exit status. **The chart does not yet provide durable report storage, missed-run alerts, or fleet policy.**
110
+
111
+ ## Safety boundary
112
+
113
+ - The generated recovery Cluster never copies `spec.plugins`, so it is not configured to archive WAL back to the source bucket. Set `recoveryObjectStore` to a separately named `ObjectStore` with read-only archive credentials. The tool compares its destination and endpoint with the source store; it cannot verify credential permissions, so test that writes are denied before relying on this boundary. Omitting this field uses the source `ObjectStore` for compatibility.
114
+ - A recovery timeout includes a machine-readable `failureReason` when the recovery Pod logs identify archive access denial or unavailable WAL. It does not copy raw logs into the report; inspect Pod and operator logs for detail.
115
+ - The drill uses the same namespace because the plugin's `ObjectStore` and its credentials are namespace-scoped. The source Cluster object is never modified.
116
+ - SQL checks run in a `BEGIN READ ONLY` transaction. Queries are limited to one `SELECT` or `WITH` statement without semicolons. Only provide trusted SQL; this is a guard against mistakes, not a security sandbox.
117
+ - By default, the tool deletes the drill Cluster even when recovery or a check fails. `retainOnFailure` leaves it for investigation and can incur storage costs. Cleanup failures turn the result red.
118
+ - The tool checks that PVCs labeled for the drill cluster disappear after Cluster deletion and fails the report if they remain. Kubernetes PersistentVolume cleanup depends on the storage class; inspect cloud volumes after a first drill.
119
+ - The planned SaaS boundary is metadata only. This repository contains no telemetry upload or hosted service.
120
+
121
+ ## Contributing and upstream distribution
122
+
123
+ Run `PYTHONPATH=src python3 -m unittest discover -s tests -v`. See [CONTRIBUTING.md](CONTRIBUTING.md) for how to add a recovery path, the live integration test gate, and the Krew package plan. We will submit only a tested, released package that stands on its own. Upstream acceptance is not assumed.
124
+
125
+ [UPSTREAM.md](UPSTREAM.md) records the specific repositories, proposed contribution value, and submission gates.
126
+
127
+ ## Recovery references
128
+
129
+ - [CloudNativePG backup guidance](https://cloudnative-pg.io/docs/devel/backup/)
130
+ - [CloudNativePG recovery documentation](https://cloudnative-pg.io/docs/devel/recovery/)
131
+ - [Barman Cloud recovery example](https://cloudnative-pg.io/plugin-barman-cloud/docs/next/concepts/)
@@ -0,0 +1,13 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ src/cnpg_drill/__init__.py
5
+ src/cnpg_drill/cli.py
6
+ src/cnpg_drill/core.py
7
+ src/cnpg_drill.egg-info/PKG-INFO
8
+ src/cnpg_drill.egg-info/SOURCES.txt
9
+ src/cnpg_drill.egg-info/dependency_links.txt
10
+ src/cnpg_drill.egg-info/entry_points.txt
11
+ src/cnpg_drill.egg-info/top_level.txt
12
+ tests/test_cli.py
13
+ tests/test_core.py
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ cnpg-drill = cnpg_drill.cli:main
@@ -0,0 +1 @@
1
+ cnpg_drill
@@ -0,0 +1,52 @@
1
+ import contextlib
2
+ import io
3
+ import json
4
+ import tempfile
5
+ import unittest
6
+ from pathlib import Path
7
+ from unittest.mock import patch
8
+
9
+ from cnpg_drill import cli
10
+
11
+ from test_core import FakeClient
12
+
13
+
14
+ class CliTest(unittest.TestCase):
15
+ def test_plan_outputs_manifest_without_creating_cluster(self):
16
+ fake = FakeClient()
17
+ with tempfile.TemporaryDirectory() as folder:
18
+ config = Path(folder) / "config.json"
19
+ config.write_text(json.dumps({"namespace": "production", "cluster": "app-db"}))
20
+ output = io.StringIO()
21
+ with patch.object(cli, "Kubectl", return_value=fake), contextlib.redirect_stdout(output):
22
+ code = cli.main(["plan", "--config", str(config)])
23
+ self.assertEqual(code, 0)
24
+ self.assertIsNone(fake.created)
25
+ self.assertEqual(json.loads(output.getvalue())["kind"], "Cluster")
26
+
27
+ def test_run_writes_machine_readable_failure_report(self):
28
+ fake = FakeClient(check_result="0")
29
+ with tempfile.TemporaryDirectory() as folder:
30
+ config = Path(folder) / "config.json"
31
+ report = Path(folder) / "report.json"
32
+ config.write_text(json.dumps({"namespace": "production", "cluster": "app-db"}))
33
+ with patch.object(cli, "Kubectl", return_value=fake), contextlib.redirect_stdout(io.StringIO()):
34
+ code = cli.main(["run", "--config", str(config), "--report", str(report)])
35
+ result = json.loads(report.read_text())
36
+ self.assertEqual(code, 1)
37
+ self.assertEqual(result["status"], "failed")
38
+ self.assertEqual(result["checks"][0]["passed"], False)
39
+
40
+ def test_run_returns_code_two_for_preflight_failure(self):
41
+ fake = FakeClient()
42
+ with tempfile.TemporaryDirectory() as folder:
43
+ config = Path(folder) / "config.json"
44
+ config.write_text(json.dumps({"namespace": "production", "cluster": "app-db", "maxBackupAgeSeconds": 1}))
45
+ with patch.object(cli, "Kubectl", return_value=fake), contextlib.redirect_stdout(io.StringIO()):
46
+ code = cli.main(["run", "--config", str(config)])
47
+ self.assertEqual(code, 2)
48
+ self.assertIsNone(fake.created)
49
+
50
+
51
+ if __name__ == "__main__":
52
+ unittest.main()
@@ -0,0 +1,267 @@
1
+ import copy
2
+ import datetime as dt
3
+ import hashlib
4
+ import json
5
+ import unittest
6
+
7
+ from cnpg_drill.core import Config, DrillError, Kubectl, build_manifest, choose_backup, prepare, recovery_failure_reason, run_drill
8
+
9
+
10
+ SOURCE = {
11
+ "metadata": {"name": "app-db", "namespace": "production", "uid": "source-uid-123"},
12
+ "spec": {
13
+ "instances": 3,
14
+ "imageName": "ghcr.io/cloudnative-pg/postgresql:17",
15
+ "storage": {"size": "10Gi", "storageClass": "fast"},
16
+ "walStorage": {"size": "2Gi"},
17
+ "plugins": [{"name": "barman-cloud.cloudnative-pg.io", "enabled": True, "isWALArchiver": True, "parameters": {"barmanObjectName": "app-store"}}],
18
+ },
19
+ }
20
+
21
+
22
+ class FakeClient:
23
+ def __init__(self, *, ready=True, check_result="1", cleanup_error=False, create_error=False, remaining_pvcs=None):
24
+ self.ready = ready
25
+ self.check_result = check_result
26
+ self.cleanup_error = cleanup_error
27
+ self.create_error = create_error
28
+ self.created = None
29
+ self.deleted = None
30
+ self.queries = []
31
+ self.remaining_pvcs = remaining_pvcs or []
32
+
33
+ def get(self, resource, name, namespace):
34
+ if name == "app-db":
35
+ return copy.deepcopy(SOURCE)
36
+ if name in ("app-store", "recovery-store", "wrong-store"):
37
+ destination = "s3://other/" if name == "wrong-store" else "s3://backups/"
38
+ return {"metadata": {"name": name}, "spec": {"configuration": {"destinationPath": destination, "endpointURL": "https://s3.example.test"}}}
39
+ return {"status": {"readyInstances": 1 if self.ready else 0, "currentPrimary": f"{name}-1" if self.ready else ""}}
40
+
41
+ def create(self, manifest, namespace):
42
+ self.created = manifest
43
+ if self.create_error:
44
+ raise DrillError("create response lost")
45
+
46
+ def delete(self, resource, name, namespace):
47
+ if self.cleanup_error:
48
+ raise DrillError("delete denied")
49
+ self.deleted = name
50
+
51
+ def exec_query(self, pod, namespace, query, database="postgres"):
52
+ self.queries.append((pod, database, query))
53
+ return self.check_result
54
+
55
+ def list_pvcs(self, cluster, namespace):
56
+ return self.remaining_pvcs
57
+
58
+ def recovery_logs(self, cluster, namespace):
59
+ return []
60
+
61
+ def list_backups(self, namespace):
62
+ stopped = (dt.datetime.now(dt.timezone.utc) - dt.timedelta(hours=1)).isoformat()
63
+ return [{"metadata": {"name": "app-db-backup"}, "spec": {"cluster": {"name": "app-db"}, "method": "plugin", "pluginConfiguration": {"name": "barman-cloud.cloudnative-pg.io"}}, "status": {"phase": "completed", "backupId": "20260929T120000", "stoppedAt": stopped, "method": "plugin"}}]
64
+
65
+
66
+ class ConfigTest(unittest.TestCase):
67
+ def test_recovery_store_name_is_validated(self):
68
+ with self.assertRaisesRegex(DrillError, "recoveryObjectStore"):
69
+ Config.from_dict({"namespace": "production", "cluster": "app-db", "recoveryObjectStore": "Other/namespace"})
70
+
71
+ def test_rejects_mutating_or_multiple_statements(self):
72
+ for query in ("DELETE FROM users", "SELECT 1; DROP TABLE users", " INSERT INTO x VALUES (1)"):
73
+ with self.subTest(query=query), self.assertRaises(DrillError):
74
+ Config.from_dict({"namespace": "production", "cluster": "app-db", "checks": [{"name": "unsafe", "query": query, "expected": "1"}]})
75
+
76
+ def test_requires_timezone_for_pitr(self):
77
+ with self.assertRaises(DrillError):
78
+ Config.from_dict({"namespace": "production", "cluster": "app-db", "targetTime": "2026-01-01T00:00:00"})
79
+
80
+ def test_database_name_rejects_connection_strings(self):
81
+ for database in ("host=elsewhere", "postgresql://elsewhere/db", "-h", "app db"):
82
+ with self.subTest(database=database), self.assertRaisesRegex(DrillError, "database"):
83
+ Config.from_dict({"namespace": "production", "cluster": "app-db", "checks": [{"name": "app-data", "database": database, "query": "SELECT 1", "expected": "1"}]})
84
+
85
+ def test_exec_uses_selected_database_in_read_only_transaction(self):
86
+ calls = []
87
+
88
+ def invoke(argv, **kwargs):
89
+ calls.append(argv)
90
+ return type("Result", (), {"returncode": 0, "stdout": "1\n", "stderr": ""})()
91
+
92
+ observed = Kubectl(invoke=invoke).exec_query("drill-1", "production", "SELECT 1", "app")
93
+ self.assertEqual(observed, "1")
94
+ self.assertEqual(calls[0][calls[0].index("-d") + 1], "app")
95
+ self.assertIn("BEGIN READ ONLY; SELECT 1; COMMIT;", calls[0])
96
+
97
+
98
+ class ManifestTest(unittest.TestCase):
99
+ def setUp(self):
100
+ self.config = Config.from_dict({"namespace": "production", "cluster": "app-db", "targetTime": "2026-01-01T00:00:00Z"})
101
+
102
+ def test_recovery_uses_source_archive_but_does_not_archive(self):
103
+ manifest = build_manifest(SOURCE, self.config, name="drill-app-db-test")
104
+ spec = manifest["spec"]
105
+ self.assertNotIn("plugins", spec)
106
+ self.assertNotIn("backup", spec)
107
+ self.assertEqual(spec["bootstrap"]["recovery"]["recoveryTarget"]["targetTime"], "2026-01-01T00:00:00Z")
108
+ self.assertEqual(spec["externalClusters"][0]["plugin"]["parameters"], {"barmanObjectName": "app-store", "serverName": "app-db"})
109
+ self.assertEqual(spec["storage"]["size"], "10Gi")
110
+ self.assertEqual(spec["walStorage"]["size"], "2Gi")
111
+
112
+ def test_separate_recovery_store_preserves_source_backup_selection(self):
113
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "recoveryObjectStore": "recovery-store"})
114
+ manifest = prepare(FakeClient(), config, name="drill-app-db-test")
115
+ self.assertEqual(manifest["spec"]["externalClusters"][0]["plugin"]["parameters"]["barmanObjectName"], "recovery-store")
116
+ self.assertEqual(manifest["metadata"]["annotations"]["cnpg-drill.dev/backup-name"], "app-db-backup")
117
+ self.assertNotIn("plugins", manifest["spec"])
118
+
119
+ def test_rejects_recovery_store_pointing_elsewhere(self):
120
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "recoveryObjectStore": "wrong-store"})
121
+ with self.assertRaisesRegex(DrillError, "destinationPath must match"):
122
+ prepare(FakeClient(), config, name="drill-app-db-test")
123
+
124
+ def test_fails_closed_on_tablespaces(self):
125
+ source = copy.deepcopy(SOURCE)
126
+ source["spec"]["tablespaces"] = [{"name": "data"}]
127
+ with self.assertRaises(DrillError):
128
+ build_manifest(source, self.config, name="drill-app-db-test")
129
+
130
+ def test_preserves_custom_application_database_identity(self):
131
+ source = copy.deepcopy(SOURCE)
132
+ source["spec"]["bootstrap"] = {"initdb": {"database": "orders", "owner": "orders_user", "secret": {"name": "orders-owner"}}}
133
+ recovery = build_manifest(source, self.config, name="drill-app-db-test")["spec"]["bootstrap"]["recovery"]
134
+ self.assertEqual(recovery["database"], "orders")
135
+ self.assertEqual(recovery["owner"], "orders_user")
136
+ self.assertEqual(recovery["secret"], {"name": "orders-owner"})
137
+
138
+ def test_copies_s3_compatibility_env_without_writer_credentials(self):
139
+ source = copy.deepcopy(SOURCE)
140
+ source["spec"]["env"] = [
141
+ {"name": "AWS_REQUEST_CHECKSUM_CALCULATION", "value": "when_required"},
142
+ {"name": "AWS_RESPONSE_CHECKSUM_VALIDATION", "value": "when_required"},
143
+ {"name": "AWS_ACCESS_KEY_ID", "valueFrom": {"secretKeyRef": {"name": "writer", "key": "access"}}},
144
+ {"name": "AWS_NO_CHUNKED_ENCODING", "valueFrom": {"secretKeyRef": {"name": "writer", "key": "flag"}}},
145
+ ]
146
+ env = build_manifest(source, self.config, name="drill-app-db-test")["spec"]["env"]
147
+ self.assertEqual(env, [
148
+ {"name": "AWS_REQUEST_CHECKSUM_CALCULATION", "value": "when_required"},
149
+ {"name": "AWS_RESPONSE_CHECKSUM_VALIDATION", "value": "when_required"},
150
+ ])
151
+
152
+
153
+ class BackupSelectionTest(unittest.TestCase):
154
+ def setUp(self):
155
+ self.config = Config.from_dict({"namespace": "production", "cluster": "app-db"})
156
+
157
+ def test_chooses_latest_completed_plugin_backup(self):
158
+ fake = FakeClient()
159
+ eligible = fake.list_backups("production")[0]
160
+ older = copy.deepcopy(eligible)
161
+ older["status"]["backupId"] = "older"
162
+ older["status"]["stoppedAt"] = (dt.datetime.now(dt.timezone.utc) - dt.timedelta(days=1)).isoformat()
163
+ failed = copy.deepcopy(eligible)
164
+ failed["status"]["phase"] = "failed"
165
+ chosen, age = choose_backup([older, failed, eligible], self.config)
166
+ self.assertEqual(chosen["status"]["backupId"], "20260929T120000")
167
+ self.assertLess(age, 7200)
168
+
169
+ def test_rejects_stale_backup(self):
170
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "maxBackupAgeSeconds": 30})
171
+ with self.assertRaisesRegex(DrillError, "above maxBackupAgeSeconds"):
172
+ choose_backup(FakeClient().list_backups("production"), config)
173
+
174
+ def test_pitr_requires_backup_before_target(self):
175
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "targetTime": "2020-01-01T00:00:00Z"})
176
+ with self.assertRaisesRegex(DrillError, "No completed"):
177
+ choose_backup(FakeClient().list_backups("production"), config)
178
+
179
+
180
+ class RunTest(unittest.TestCase):
181
+ def setUp(self):
182
+ self.config = Config.from_dict({"namespace": "production", "cluster": "app-db"})
183
+
184
+ def test_success_reports_check_and_cleans_up(self):
185
+ client = FakeClient()
186
+ report = run_drill(client, self.config)
187
+ self.assertEqual(report["status"], "passed")
188
+ self.assertEqual(report["backupID"], "20260929T120000")
189
+ self.assertLess(report["backupAgeSeconds"], 7200)
190
+ self.assertEqual(report["cleanup"], "cluster-and-pvcs-deleted")
191
+ self.assertEqual(client.deleted, report["drillCluster"])
192
+ self.assertEqual(report["checks"], [{"name": "connection", "database": "postgres", "passed": True, "observedSha256": hashlib.sha256(b"1").hexdigest()}])
193
+
194
+ def test_check_can_query_application_database(self):
195
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "checks": [{"name": "app-data", "database": "app", "query": "SELECT 1", "expected": "1"}]})
196
+ client = FakeClient()
197
+ report = run_drill(client, config)
198
+ self.assertEqual(report["status"], "passed")
199
+ self.assertEqual(client.queries[0][1], "app")
200
+ self.assertEqual(report["checks"][0]["database"], "app")
201
+
202
+ def test_failed_check_still_cleans_up(self):
203
+ client = FakeClient(check_result="0")
204
+ report = run_drill(client, self.config)
205
+ self.assertEqual(report["status"], "failed")
206
+ self.assertEqual(report["cleanup"], "cluster-and-pvcs-deleted")
207
+ self.assertNotIn("got '0'", report["error"])
208
+
209
+ def test_cleanup_failure_fails_drill(self):
210
+ client = FakeClient(cleanup_error=True)
211
+ report = run_drill(client, self.config)
212
+ self.assertEqual(report["status"], "failed")
213
+ self.assertEqual(report["cleanup"], "failed")
214
+
215
+ def test_attempts_cleanup_after_ambiguous_create_failure(self):
216
+ client = FakeClient(create_error=True)
217
+ report = run_drill(client, self.config)
218
+ self.assertEqual(report["status"], "failed")
219
+ self.assertEqual(client.deleted, client.created["metadata"]["name"])
220
+ self.assertNotEqual(report.get("phase"), "preflight")
221
+
222
+ def test_stale_backup_is_preflight_failure(self):
223
+ client = FakeClient()
224
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "maxBackupAgeSeconds": 1})
225
+ report = run_drill(client, config)
226
+ self.assertEqual(report["phase"], "preflight")
227
+ self.assertIsNone(client.created)
228
+
229
+ def test_retention_only_on_failure(self):
230
+ client = FakeClient(check_result="0")
231
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "retainOnFailure": True})
232
+ report = run_drill(client, config)
233
+ self.assertEqual(report["cleanup"], "retained-for-investigation")
234
+ self.assertIsNone(client.deleted)
235
+
236
+ def test_freshness_check_passes_without_logging_timestamp(self):
237
+ import datetime as dt
238
+ observed = (dt.datetime.now(dt.timezone.utc) - dt.timedelta(seconds=10)).isoformat()
239
+ client = FakeClient(check_result=observed)
240
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "checks": [{"name": "recent-order", "query": "SELECT now()", "maxAgeSeconds": 60}]})
241
+ report = run_drill(client, config)
242
+ self.assertEqual(report["status"], "passed")
243
+ self.assertNotIn(observed, str(report))
244
+ self.assertLess(report["checks"][0]["ageSeconds"], 60)
245
+
246
+ def test_recovery_diagnostic_classifies_without_leaking_logs(self):
247
+ logs = ['{"error":"WAL file 000000010000000000000005 not found in archive s3://private/"}']
248
+ self.assertEqual(recovery_failure_reason(logs), "wal_unavailable")
249
+ self.assertEqual(recovery_failure_reason(['{"error":"AccessDenied: secret-bucket"}']), "archive_access_denied")
250
+ self.assertIsNone(recovery_failure_reason(['{"msg":"restored log file from archive"}']))
251
+
252
+ def test_timeout_reports_safe_reason_and_cleans_up(self):
253
+ class FailedRecovery(FakeClient):
254
+ def recovery_logs(self, cluster, namespace):
255
+ return ['{"error":"AccessDenied for s3://secret-path/"}']
256
+
257
+ client = FailedRecovery(ready=False)
258
+ times = iter([0, 1, 31, 32, 33])
259
+ config = Config.from_dict({"namespace": "production", "cluster": "app-db", "timeoutSeconds": 30})
260
+ report = run_drill(client, config, sleep=lambda _: None, monotonic=lambda: next(times))
261
+ self.assertEqual(report["failureReason"], "archive_access_denied")
262
+ self.assertNotIn("secret-path", str(report))
263
+ self.assertEqual(report["cleanup"], "cluster-and-pvcs-deleted")
264
+
265
+
266
+ if __name__ == "__main__":
267
+ unittest.main()