devopsiq 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agent/__init__.py +5 -0
- agent/agent.py +232 -0
- agent/investigation.py +339 -0
- agent/prompts.py +125 -0
- agent/store.py +147 -0
- devopsiq-0.1.0.dist-info/METADATA +662 -0
- devopsiq-0.1.0.dist-info/RECORD +30 -0
- devopsiq-0.1.0.dist-info/WHEEL +5 -0
- devopsiq-0.1.0.dist-info/entry_points.txt +2 -0
- devopsiq-0.1.0.dist-info/licenses/LICENSE +21 -0
- devopsiq-0.1.0.dist-info/top_level.txt +3 -0
- main.py +310 -0
- tools/__init__.py +6 -0
- tools/ansible.py +110 -0
- tools/argocd.py +91 -0
- tools/base.py +112 -0
- tools/cloud.py +101 -0
- tools/docker.py +280 -0
- tools/git_ci.py +257 -0
- tools/helm.py +168 -0
- tools/investigation.py +357 -0
- tools/istio.py +43 -0
- tools/kubernetes.py +464 -0
- tools/monitoring.py +162 -0
- tools/newrelic.py +167 -0
- tools/preflight.py +59 -0
- tools/registry.py +49 -0
- tools/system.py +162 -0
- tools/terraform.py +90 -0
- tools/trivy.py +83 -0
|
@@ -0,0 +1,662 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: devopsiq
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: A from-scratch, read-only AI agent that investigates real DevOps problems — Kubernetes, Docker, Helm, Argo CD, Istio, Terraform, GitHub Actions, cloud, New Relic and more.
|
|
5
|
+
Author: DevOpsAbhii
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/DevOpsAbhii/devops-ai-agent
|
|
8
|
+
Project-URL: Repository, https://github.com/DevOpsAbhii/devops-ai-agent
|
|
9
|
+
Project-URL: Issues, https://github.com/DevOpsAbhii/devops-ai-agent/issues
|
|
10
|
+
Keywords: devops,kubernetes,docker,helm,argocd,istio,terraform,github-actions,observability,newrelic,incident-response,root-cause-analysis,ai-agent,cli
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Environment :: Console
|
|
13
|
+
Classifier: Intended Audience :: Developers
|
|
14
|
+
Classifier: Intended Audience :: System Administrators
|
|
15
|
+
Classifier: Operating System :: POSIX :: Linux
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
21
|
+
Classifier: Topic :: System :: Systems Administration
|
|
22
|
+
Classifier: Topic :: Software Development :: Debuggers
|
|
23
|
+
Requires-Python: >=3.10
|
|
24
|
+
Description-Content-Type: text/markdown
|
|
25
|
+
License-File: LICENSE
|
|
26
|
+
Requires-Dist: openai>=1.40.0
|
|
27
|
+
Requires-Dist: python-dotenv>=1.0.1
|
|
28
|
+
Dynamic: license-file
|
|
29
|
+
|
|
30
|
+
# DevOps AI Agent
|
|
31
|
+
|
|
32
|
+
An AI agent that investigates real DevOps problems. The end goal: ask it
|
|
33
|
+
something like *"Why is my Kubernetes pod in CrashLoopBackOff?"* and have it
|
|
34
|
+
gather evidence, reason about the evidence, identify the likely root cause,
|
|
35
|
+
recommend remediation, and give verification steps.
|
|
36
|
+
|
|
37
|
+
**Phase 5** broadens the agent from
|
|
38
|
+
Kubernetes-only to the wider DevOps surface: Linux systemd services and
|
|
39
|
+
journals, Docker containers, Terraform state/plan, and git/GitHub pull
|
|
40
|
+
requests — every tool still read-only, still a fixed allowlisted argv
|
|
41
|
+
template. **Phase 6 adds the automation surface:** a one-shot CLI mode
|
|
42
|
+
(`python main.py "problem"` — exit-code driven, cron/CI-friendly) and a
|
|
43
|
+
structured JSON export of the investigation report, so pipelines can act on
|
|
44
|
+
the verdict instead of parsing markdown. **Phase 7 adds persistence:** the
|
|
45
|
+
investigation record is auto-saved on every change to
|
|
46
|
+
`~/.devops-ai-agent/investigations/` (overridable with `--store-dir` or
|
|
47
|
+
`AGENT_STORE_DIR`), survives CLI exits, and the REPL resumes the newest
|
|
48
|
+
in-progress record at startup. **Phase 8 broadens the tool surface to 47
|
|
49
|
+
read-only tools:** Kubernetes depth (pod listing, `kubectl top` resource
|
|
50
|
+
usage, HPAs, PVCs, contexts), Docker depth (networks, volumes, disk usage),
|
|
51
|
+
GitHub Actions via `gh` (runs, run jobs, workflows), cloud identity
|
|
52
|
+
(AWS/GCP/Azure), monitoring/logging (Prometheus, Loki, Grafana — endpoints
|
|
53
|
+
from environment config), and Ansible listing (inventory, playbook tasks).
|
|
54
|
+
**Phase 9 adds 11 trending-market tools (58 total):** New Relic (NRQL +
|
|
55
|
+
open alerts over NerdGraph — credentials from env), Trivy image
|
|
56
|
+
vulnerability scanning, Helm releases (list/status/history), Argo CD
|
|
57
|
+
(GitOps app sync/health), Istio mesh proxy status, and Docker Compose
|
|
58
|
+
(project and service listing). The repository is git-tracked.
|
|
59
|
+
Every phase still built from scratch — no LangChain, LangGraph,
|
|
60
|
+
AutoGen, CrewAI, or MCP.
|
|
61
|
+
|
|
62
|
+
> **Want to use this in your own project?** See [INTEGRATION.md](INTEGRATION.md)
|
|
63
|
+
> for the three integration levels: drive it as a CLI from cron/CI (exit
|
|
64
|
+
> codes + `--json`), embed it as a Python library, or extend it with your
|
|
65
|
+
> own read-only tools.
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
69
|
+
## 1. What this project is
|
|
70
|
+
|
|
71
|
+
The agent is built intentionally: we own the core architecture — how the
|
|
72
|
+
model is called, how conversation history flows, how tool selection +
|
|
73
|
+
execution + evidence feedback work — instead of depending on a framework
|
|
74
|
+
for it.
|
|
75
|
+
|
|
76
|
+
In Phase 5 the agent:
|
|
77
|
+
- holds a conversation with **GLM 5.3** through **OpenRouter**;
|
|
78
|
+
- has **58 real, read-only tools** across fifteen domains: host facts,
|
|
79
|
+
Kubernetes (12 tools), Linux system (4), Docker + Compose (10),
|
|
80
|
+
Terraform (3), Helm (3), Argo CD (2), Istio (1), security/Trivy (1),
|
|
81
|
+
git/GitHub + Actions (7), cloud identity (4), monitoring/logging (3),
|
|
82
|
+
New Relic (2), Ansible (2), plus the 3 investigation tools;
|
|
83
|
+
- runs a **tool-use loop**: when the model decides a question needs evidence,
|
|
84
|
+
it requests the tool, the application executes it locally, and the real
|
|
85
|
+
output is fed back to the model, which then answers from it;
|
|
86
|
+
- runs an **investigation loop**: for a reported problem it plans a multi-step
|
|
87
|
+
investigation — open a record with initial hypotheses, gather evidence with
|
|
88
|
+
the right domain tools, update each hypothesis's verdict as evidence
|
|
89
|
+
builds, conclude with root cause + remediation + verification, and end with
|
|
90
|
+
a structured report. The tracked state lives in the application, so
|
|
91
|
+
`/investigation` and `/report` show it directly.
|
|
92
|
+
|
|
93
|
+
It is **read-only by design**: every tool is a static, allowlisted command
|
|
94
|
+
template — the model can never pass arbitrary command text, can never select
|
|
95
|
+
a mutating verb (no delete/restart/edit/apply/scale/exec/run/rm/reset/
|
|
96
|
+
destroy anywhere), and no mutating capability exists. Object and unit names
|
|
97
|
+
are validated before reaching any CLI; flags are injected only by fixed
|
|
98
|
+
templates, never by model text. The investigation tools mutate only the
|
|
99
|
+
agent's in-memory record. Any future mutating capability will only ever
|
|
100
|
+
arrive behind an explicit human-approval gate.
|
|
101
|
+
|
|
102
|
+
## 2. Current architecture
|
|
103
|
+
|
|
104
|
+
```
|
|
105
|
+
You (terminal)
|
|
106
|
+
│ plain text
|
|
107
|
+
▼
|
|
108
|
+
main.py ......................... CLI REPL, .env loading, error handling
|
|
109
|
+
│
|
|
110
|
+
▼
|
|
111
|
+
agent/agent.py (DevOpsAgent) .... system prompt + conversation history
|
|
112
|
+
│
|
|
113
|
+
▼ the tool-use loop, inside DevOpsAgent._complete():
|
|
114
|
+
│
|
|
115
|
+
│ attach tool schemas -> call model
|
|
116
|
+
│ │
|
|
117
|
+
│ ├─ model requests tool(s)? ── yes ─▶ tools/registry.execute_tool()
|
|
118
|
+
│ │ │ per call, in order
|
|
119
|
+
│ │ │ └─ allowlisted read-only
|
|
120
|
+
│ │ │ command runs for real
|
|
121
|
+
│ │ │ └─ result echoed back as
|
|
122
|
+
│ │ │ a "tool" message
|
|
123
|
+
│ │ ◀── loop calls the model again
|
|
124
|
+
│ └─ model answers in plain text? ── yes ─▶ done
|
|
125
|
+
│ (safety cap: max 10 tool-use turns, then abort)
|
|
126
|
+
▼
|
|
127
|
+
OpenRouter (https://openrouter.ai/api/v1)
|
|
128
|
+
▼
|
|
129
|
+
GLM 5.3 (z-ai/glm-5.3)
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
Module map:
|
|
133
|
+
|
|
134
|
+
| Path | Responsibility |
|
|
135
|
+
| --------------------------- | ------------------------------------------------------------ |
|
|
136
|
+
| `main.py` | REPL loop, one-shot CLI, environment loading, `/investigate` commands, error messages |
|
|
137
|
+
| `agent/agent.py` | `DevOpsAgent` — client, history, `ask()`, `_complete()` loop |
|
|
138
|
+
| `agent/prompts.py` | The system prompt (versioned/tested separately) |
|
|
139
|
+
| `agent/investigation.py` | First-class investigation record: hypotheses, verdicts, evidence, report renderer (pure data) |
|
|
140
|
+
| `agent/store.py` | `InvestigationStore` — one JSON file per record, atomic writes, resume/list (Phase 7) |
|
|
141
|
+
| `tools/base.py` | Tool contract: `Tool`, `ToolError`, `read_command_output()` |
|
|
142
|
+
| `tools/registry.py` | `register/get_tools/execute_tool` — tools declared & executed here |
|
|
143
|
+
| `tools/preflight.py` | `system_info` — host facts (allowlisted read-only commands) |
|
|
144
|
+
| `tools/kubernetes.py` | 12 kubectl tools: pods, pod status/logs, deployments, events, nodes, top (cpu/mem), hpa, pvc, services, contexts |
|
|
145
|
+
| `tools/system.py` | systemd/journal/ss/ps tools — services, journals, ports, top processes |
|
|
146
|
+
| `tools/docker.py` | read-only docker tools — ps, inspect, logs, stats, images, networks, volumes, disk usage, compose ls/ps |
|
|
147
|
+
| `tools/terraform.py` | `tf_show` / `tf_state_list` / `tf_plan` (working directory) |
|
|
148
|
+
| `tools/helm.py` | `helm_list` / `helm_status` / `helm_history` — release reads only (Phase 9) |
|
|
149
|
+
| `tools/argocd.py` | `argocd_apps` / `argocd_app_status` — GitOps app reads (Phase 9) |
|
|
150
|
+
| `tools/istio.py` | `istioctl_proxy_status` — mesh sync view, no arguments (Phase 9) |
|
|
151
|
+
| `tools/trivy.py` | `trivy_image_scan` — vulnerability report for one image (Phase 9) |
|
|
152
|
+
| `tools/newrelic.py` | `newrelic_nrql` / `newrelic_alerts` — NerdGraph over curl, credentials from env only (Phase 9) |
|
|
153
|
+
| `tools/git_ci.py` | git status/log/diff + `gh_prs`, `gh_runs`, `gh_run_view`, `gh_workflows` (working directory) |
|
|
154
|
+
| `tools/cloud.py` | read-only cloud identity/listing: `aws_identity`, `gcloud_identity`, `az_account`, `az_groups` |
|
|
155
|
+
| `tools/monitoring.py` | `prom_query`, `loki_query`, `grafana_health` — endpoints from env config only |
|
|
156
|
+
| `tools/ansible.py` | `ansible_inventory`, `ansible_playbook_tasks` — listing modes only |
|
|
157
|
+
| `tools/investigation.py` | `investigation_begin` / `investigation_record` / `investigation_conclude` — meta-tools for the investigation record |
|
|
158
|
+
| `tests/test_phase2.py` | Offline suite: contract, safety, loop (stdlib `unittest`) |
|
|
159
|
+
| `tests/test_phase3.py` | Offline suite: k8s command lines + verb-allowlist (fake kubectl) |
|
|
160
|
+
| `tests/test_phase4.py` | Offline suite: investigation state, meta-tools, loop-driven investigation |
|
|
161
|
+
| `tests/test_phase5.py` | Offline suite: argv templates + name validation for all new domains (fake CLIs) |
|
|
162
|
+
| `tests/test_automation.py` | Offline suite: JSON reports + one-shot CLI (Phase 6) |
|
|
163
|
+
| `tests/test_phase7.py` | Offline suite: round-trip serialization, store files, auto-save, resume, CLI flags (Phase 7) |
|
|
164
|
+
| `tests/test_phase8.py` | Offline suite: argv templates + validation for the 21 Phase 8 tools (fake CLIs, env-based monitoring) |
|
|
165
|
+
| `tests/test_phase9.py` | Offline suite: New Relic env-credential + payload tests, trivy/helm/argocd/istio/compose argv templates (Phase 9) |
|
|
166
|
+
|
|
167
|
+
### The tool-use loop
|
|
168
|
+
|
|
169
|
+
All model calls funnel through `DevOpsAgent._complete()`. Each turn:
|
|
170
|
+
|
|
171
|
+
1. The full history is sent with every tool's schema attached (`tools=`).
|
|
172
|
+
2. If the reply contains tool calls, the assistant tool-request turn is
|
|
173
|
+
echoed into history verbatim, and each call is executed locally via
|
|
174
|
+
`tools/registry.execute_tool(name, arguments)`.
|
|
175
|
+
3. Every real result is appended as a sibling `{"role": "tool"}` message
|
|
176
|
+
pinned to its `tool_call_id`.
|
|
177
|
+
4. The loop calls the model again — giving it the evidence to reflect on —
|
|
178
|
+
and repeats until it answers in plain text.
|
|
179
|
+
5. `MAX_TOOL_ITERATIONS = 10` aborts runaway loops.
|
|
180
|
+
|
|
181
|
+
### The investigation loop (Phase 4)
|
|
182
|
+
|
|
183
|
+
For a reported problem (a pod in CrashLoopBackOff, a broken rollout) the
|
|
184
|
+
same loop carries a *plan* as well as evidence:
|
|
185
|
+
|
|
186
|
+
1. The model opens a record with `investigation_begin` (one-line problem +
|
|
187
|
+
2–4 initial hypotheses, which become H1, H2, …).
|
|
188
|
+
2. It plans what it needs, then calls the read-only tools one deliberate
|
|
189
|
+
step at a time.
|
|
190
|
+
3. Every finding is recorded with `investigation_record`: evidence notes
|
|
191
|
+
(linked to a hypothesis when they bear on one) and verdicts
|
|
192
|
+
(supported / refuted / confirmed).
|
|
193
|
+
4. Each record call returns the **live tracker** (`Problem`, status,
|
|
194
|
+
hypotheses with verdicts, evidence), so the model always knows where it
|
|
195
|
+
stands without inspecting all of history.
|
|
196
|
+
5. When evidence is sufficient the model calls `investigation_conclude`
|
|
197
|
+
with root cause, remediation *recommendations* (nothing is ever
|
|
198
|
+
executed), verification steps, and confidence — then ends its answer
|
|
199
|
+
with the structured report.
|
|
200
|
+
|
|
201
|
+
The record lives in the **application**, not just the conversation: the
|
|
202
|
+
report is rendered deterministically from it (agent/investigation.py), and
|
|
203
|
+
the CLI exposes it directly:
|
|
204
|
+
|
|
205
|
+
| Command | What it does |
|
|
206
|
+
| ------------------------ | --------------------------------------------------- |
|
|
207
|
+
| `/investigate <problem>` | Open a formal investigation (same path the model uses) |
|
|
208
|
+
| `/investigation` | Show the live tracked state (hypotheses/evidence) |
|
|
209
|
+
| `/investigations` | List saved records on disk (newest first; `← active` marks the live one) |
|
|
210
|
+
| `/report` | Show the canonical report (once concluded) |
|
|
211
|
+
| `/endinvestigation` | Clear the record (memory only — the saved copy stays as history) |
|
|
212
|
+
|
|
213
|
+
The report the agent ends with and `/report` render are kept consistent by
|
|
214
|
+
construction: `tool` results confirm each record call and the tracker, and
|
|
215
|
+
`render_report()` is the single report format.
|
|
216
|
+
|
|
217
|
+
Example (live, against a real cluster):
|
|
218
|
+
|
|
219
|
+
```
|
|
220
|
+
You: /investigate Pod crashloop-6f8c… is in CrashLoopBackOff in default ns.
|
|
221
|
+
(…) investigate.
|
|
222
|
+
Agent: (investigation_begin → k8s_pod_status → k8s_pod_logs →
|
|
223
|
+
k8s_deployment_status → verdicts on H1/H2 → investigation_conclude)
|
|
224
|
+
ends with: Facts → Hypotheses → Root cause → Remediation (recommended)
|
|
225
|
+
→ Verification steps.
|
|
226
|
+
|
|
227
|
+
You: /report
|
|
228
|
+
Agent: # Investigation report … (canonical, rendered from the record)
|
|
229
|
+
```
|
|
230
|
+
|
|
231
|
+
### The tools (Phases 5, 8 and 9)
|
|
232
|
+
|
|
233
|
+
Every tool runs a fixed, allowlisted command — the model only picks
|
|
234
|
+
arguments (names, counts, namespaces), never command text.
|
|
235
|
+
|
|
236
|
+
| Domain | Tool | Backing command (read-only) |
|
|
237
|
+
| ---------------- | --------------------- | ---------------------------------------------------- |
|
|
238
|
+
| Host | `system_info` | `date`/`uname`/`uptime`/`df`/`free` |
|
|
239
|
+
| Kubernetes | `k8s_pod_status` | `kubectl get pod <pod> -n <ns> -o json` |
|
|
240
|
+
| | `k8s_pod_logs` | `kubectl logs <pod> -n <ns> --tail=<n>` |
|
|
241
|
+
| | `k8s_deployment_status` | `kubectl get deployment <dep> -n <ns> -o json` |
|
|
242
|
+
| | `k8s_events` | `kubectl get events -n <ns> --sort-by=.lastTimestamp -o wide [--field-selector involvedObject.name=<obj>]` |
|
|
243
|
+
| | `k8s_nodes` | `kubectl get nodes -o json` |
|
|
244
|
+
| | `k8s_services` | `kubectl get services -n <ns> -o json` |
|
|
245
|
+
| | `k8s_pods` | `kubectl get pods -n <ns> -o wide` |
|
|
246
|
+
| | `k8s_top_pods` | `kubectl top pods -n <ns> [--sort-by=cpu\|memory]` |
|
|
247
|
+
| | `k8s_top_nodes` | `kubectl top nodes` |
|
|
248
|
+
| | `k8s_hpa` | `kubectl get hpa [<name>] -n <ns> -o json` |
|
|
249
|
+
| | `k8s_pvc` | `kubectl get pvc [<name>] -n <ns> -o json` |
|
|
250
|
+
| | `k8s_contexts` | `kubectl config get-contexts` (listing only — never `use-context`) |
|
|
251
|
+
| Linux system | `sys_service_status` | `systemctl status <unit> --no-pager` |
|
|
252
|
+
| | `sys_service_logs` | `journalctl -u <unit> --no-pager -n <n>` |
|
|
253
|
+
| | `sys_open_ports` | `ss -tlnp` |
|
|
254
|
+
| | `sys_top_processes` | `ps aux --sort=-%cpu --no-headers` |
|
|
255
|
+
| Docker | `docker_ps` | `docker ps -a` |
|
|
256
|
+
| | `docker_inspect` | `docker inspect <name>` |
|
|
257
|
+
| | `docker_logs` | `docker logs --tail <n> <name>` |
|
|
258
|
+
| | `docker_stats` | `docker stats --no-stream` (the flag is pinned — see below) |
|
|
259
|
+
| | `docker_images` | `docker images` |
|
|
260
|
+
| | `docker_networks` | `docker network ls` |
|
|
261
|
+
| | `docker_volumes` | `docker volume ls` |
|
|
262
|
+
| | `docker_disk_usage` | `docker system df` |
|
|
263
|
+
| | `docker_compose_ls` | `docker compose ls` |
|
|
264
|
+
| | `docker_compose_ps` | `docker compose [-p <project>] ps -a` |
|
|
265
|
+
| Helm | `helm_list` | `helm list -n <ns>` / `helm list --all-namespaces` |
|
|
266
|
+
| | `helm_status` | `helm status <release> -n <ns>` |
|
|
267
|
+
| | `helm_history` | `helm history <release> -n <ns> --max <n>` |
|
|
268
|
+
| Argo CD | `argocd_apps` | `argocd app list --output json` |
|
|
269
|
+
| | `argocd_app_status` | `argocd app get <app>` |
|
|
270
|
+
| Istio | `istioctl_proxy_status` | `istioctl proxy-status` (no arguments) |
|
|
271
|
+
| Security | `trivy_image_scan` | `trivy image --scanners vuln --format table <image>` |
|
|
272
|
+
| New Relic | `newrelic_nrql` | `curl -H "API-Key: …" -d <json payload> https://api.newrelic.com/graphql` |
|
|
273
|
+
| | `newrelic_alerts` | same pinned curl + payload (no model input) |
|
|
274
|
+
| Terraform | `tf_show` | `terraform show -no-color` |
|
|
275
|
+
| | `tf_state_list` | `terraform state list` |
|
|
276
|
+
| | `tf_plan` | `terraform plan -no-color -input=false` |
|
|
277
|
+
| git / GitHub | `git_repo_status` | `git status --short --branch` |
|
|
278
|
+
| | `git_log` | `git log --oneline -n <n>` |
|
|
279
|
+
| | `git_diff` | `git diff --stat HEAD` |
|
|
280
|
+
| | `gh_prs` | `gh pr list --limit <n> --json number,title,state,...` |
|
|
281
|
+
| | `gh_runs` | `gh run list --limit <n> --json databaseId,displayTitle,status,...` |
|
|
282
|
+
| | `gh_run_view` | `gh run view <id> --json status,conclusion,jobs` (id digits-only) |
|
|
283
|
+
| | `gh_workflows` | `gh workflow list --limit <n> --json id,name,state` |
|
|
284
|
+
| Cloud identity | `aws_identity` | `aws sts get-caller-identity --output json` |
|
|
285
|
+
| | `gcloud_identity` | `gcloud config list --format=json` |
|
|
286
|
+
| | `az_account` | `az account show` |
|
|
287
|
+
| | `az_groups` | `az group list` |
|
|
288
|
+
| Monitoring | `prom_query` | `curl <PROMETHEUS_URL>/api/v1/query?query=<urlencoded PromQL>` |
|
|
289
|
+
| | `loki_query` | `curl <LOKI_URL>/loki/api/v1/query?query=<urlencoded LogQL>&limit=<n>` |
|
|
290
|
+
| | `grafana_health` | `curl <GRAFANA_URL>/api/health` |
|
|
291
|
+
| Ansible | `ansible_inventory` | `ansible-inventory [--inventory <source>] --list` |
|
|
292
|
+
| | `ansible_playbook_tasks` | `ansible-playbook --list-tasks --list-hosts <playbook>` |
|
|
293
|
+
| Investigation | `investigation_begin` / `investigation_record` / `investigation_conclude` | record only — memory, no external command |
|
|
294
|
+
|
|
295
|
+
Notes:
|
|
296
|
+
|
|
297
|
+
- Output is the **real** CLI output — Python does not re-parse it; the model
|
|
298
|
+
reads exactly what an engineer would see. Verbatim quoting is a hard
|
|
299
|
+
prompt rule.
|
|
300
|
+
- `docker stats` is **always** `--no-stream` (the pinned flag is what keeps
|
|
301
|
+
it from following forever and hanging the turn).
|
|
302
|
+
- `terraform plan` is a **dry run** — it computes the diff, mutates nothing;
|
|
303
|
+
`-input=false` keeps it from ever prompting. Requires an initialized
|
|
304
|
+
directory (`terraform init`); on a fresh dir the model reports terraform's
|
|
305
|
+
own error, honestly.
|
|
306
|
+
- Terraform and git/gh tools operate on the **current working directory** the
|
|
307
|
+
agent was launched from (no path parameters — that keeps traversal out).
|
|
308
|
+
- `gh_prs` needs the `gh` CLI authenticated and a GitHub remote; otherwise
|
|
309
|
+
the exact CLI error is returned.
|
|
310
|
+
- `kubectl top` needs **metrics-server** in the cluster; where it is absent,
|
|
311
|
+
kubectl's exact error is returned (that is honest behavior, not a bug).
|
|
312
|
+
- `k8s_contexts` is a listing of kubeconfig contexts only — the agent can
|
|
313
|
+
never run `config use-context` and change which cluster is in use.
|
|
314
|
+
- **Monitoring endpoints come from the environment, never the model:**
|
|
315
|
+
`prom_query`/`loki_query`/`grafana_health` read `PROMETHEUS_URL` /
|
|
316
|
+
`LOKI_URL` / `GRAFANA_URL`; unset or non-http(s) values are honest
|
|
317
|
+
ToolErrors naming the variable. The model supplies only the PromQL/LogQL
|
|
318
|
+
text, which is percent-encoded into the query string; the pinned curl argv
|
|
319
|
+
(`--proto =https,http`, GET only) blocks `file://` and everything but the
|
|
320
|
+
configured host.
|
|
321
|
+
- Cloud tools are identity/account level only and take no arguments at all;
|
|
322
|
+
Ansible tools are listing modes only (`--list`, `--list-tasks`) with
|
|
323
|
+
relative-path validation (no `/`, no leading `-`, no `..`).
|
|
324
|
+
- **Helm tools only read** (list/status/history — never install/upgrade/
|
|
325
|
+
rollback/uninstall); **Argo CD tools only list/get** (never sync/rollback/
|
|
326
|
+
delete); `istioctl_proxy_status` takes no arguments at all;
|
|
327
|
+
`trivy_image_scan` validates the image reference (no leading `-`, no
|
|
328
|
+
spaces) and warns its first run may take minutes (CVE DB download).
|
|
329
|
+
- **New Relic credentials are environment-configured, never model-chosen:**
|
|
330
|
+
`NEW_RELIC_API_KEY` (passed only as a curl header value) and
|
|
331
|
+
`NEW_RELIC_ACCOUNT_ID` (digits only). The NRQL text travels inside a
|
|
332
|
+
`json.dumps`-built payload as a GraphQL *variable*, so it can never
|
|
333
|
+
escape its string slot; only NerdGraph queries are ever sent, never
|
|
334
|
+
mutations.
|
|
335
|
+
- Log tails (`--tail`, `-n`, `--limit`) are bounded integers, validated
|
|
336
|
+
1–500 (git log 1–100, gh 1–50, Loki limit 1–1000, helm history 1–50).
|
|
337
|
+
- Any missing CLI / unreachable target returns the exact error — nothing is
|
|
338
|
+
invented (verified live for kubectl, systemctl, journalctl, docker,
|
|
339
|
+
terraform, git).
|
|
340
|
+
|
|
341
|
+
### How read-only is enforced (defense in depth)
|
|
342
|
+
|
|
343
|
+
- The tool schemas only allow picking names/counts/namespaces from validated
|
|
344
|
+
arguments — there is no way to pass command text to any CLI.
|
|
345
|
+
- Every tool is a fixed argv template. No `sh -c` anywhere, so nothing is
|
|
346
|
+
ever parsed by a shell; flags the model might abuse (e.g. `--no-stream`
|
|
347
|
+
for docker stats, `-input=false` for terraform plan) are hard-coded into
|
|
348
|
+
the template and cannot be removed or added.
|
|
349
|
+
- Only read-only verbs exist per domain — `get`/`logs`/`top` and
|
|
350
|
+
`config get-contexts` (kubectl), `status`/`-u`/`ss`/`ps` (system),
|
|
351
|
+
`ps`/`inspect`/`logs`/`stats`/`images`/`network ls`/`volume ls`/
|
|
352
|
+
`system df`/`compose ls`/`compose ps` (docker), `show`/`state list`/
|
|
353
|
+
`plan` (terraform), `list`/`status`/`history` (helm), `app list`/
|
|
354
|
+
`app get` (argocd), `proxy-status` (istioctl), `image --scanners vuln`
|
|
355
|
+
(trivy), `status`/`log`/`diff`/`pr list`/`run list`/`run view`/
|
|
356
|
+
`workflow list` (git/gh), `--list`/`--list-tasks` (ansible), GET-only
|
|
357
|
+
curl with a pinned argv (monitoring), query-only NerdGraph with a pinned
|
|
358
|
+
curl (New Relic). No delete, restart, edit, apply, scale, exec, run, rm,
|
|
359
|
+
pull, push, commit, reset, merge, use-context, playbook-run, install,
|
|
360
|
+
upgrade, rollback, uninstall, sync, destroy — by construction.
|
|
361
|
+
- Names are validated per domain before reaching any CLI: Kubernetes object
|
|
362
|
+
names (DNS style), systemd unit names (no `/`, no leading `-`), Docker
|
|
363
|
+
names (no `/`), GitHub run ids (digits only), Ansible sources (relative
|
|
364
|
+
paths, no `..`), and no path parameters at all for terraform/git/gh tools.
|
|
365
|
+
This blocks flag and path injection.
|
|
366
|
+
- The executor re-validates every argument. Never trust the model.
|
|
367
|
+
- `--request-timeout`/timeouts bound slow or hanging commands, and logs are
|
|
368
|
+
tail-bounded.
|
|
369
|
+
- Unrecognized tools/arguments return `Tool error: ...` to the model instead
|
|
370
|
+
of executing.
|
|
371
|
+
- Tool output is truncated at 8,000 characters per result.
|
|
372
|
+
|
|
373
|
+
## 3. Why GLM 5.3
|
|
374
|
+
|
|
375
|
+
- GLM 5.3 (`z-ai/glm-5.3` on OpenRouter) is a capable, cost-effective
|
|
376
|
+
general model — with strong instruction-following and **function-calling**
|
|
377
|
+
support, which the tool loop has demonstrated live in every phase.
|
|
378
|
+
- Served by Zhipu AI through a single OpenRouter endpoint, so there is no
|
|
379
|
+
separate vendor API to manage.
|
|
380
|
+
- An investigation loop sends many tokens (system prompt, tool schemas, tool
|
|
381
|
+
output, history); GLM 5.3 keeps that cost sustainable.
|
|
382
|
+
|
|
383
|
+
## 4. How OpenRouter fits
|
|
384
|
+
|
|
385
|
+
[OpenRouter](https://openrouter.ai) is a model gateway: one API key, one
|
|
386
|
+
OpenAI-compatible endpoint, access to many models. We call
|
|
387
|
+
|
|
388
|
+
```
|
|
389
|
+
POST https://openrouter.ai/api/v1/chat/completions
|
|
390
|
+
```
|
|
391
|
+
|
|
392
|
+
with model `z-ai/glm-5.3` and standard `tools` / `tool` messages. The
|
|
393
|
+
official `openai` Python SDK works as our client with two config lines:
|
|
394
|
+
|
|
395
|
+
```python
|
|
396
|
+
OpenAI(api_key=..., base_url="https://openrouter.ai/api/v1")
|
|
397
|
+
```
|
|
398
|
+
|
|
399
|
+
Swapping to another OpenRouter model later is a one-line change (an env var
|
|
400
|
+
today); moving to any other OpenAI-compatible provider changes only
|
|
401
|
+
`agent/agent.py` configuration.
|
|
402
|
+
|
|
403
|
+
## 5. Installation
|
|
404
|
+
|
|
405
|
+
Three ways — pick one. Whatever you choose, the agent also needs the CLIs of
|
|
406
|
+
the domains you use (kubectl, docker, gh, … — each tool reports its exact
|
|
407
|
+
error if its CLI is missing, so a partial install is fine).
|
|
408
|
+
|
|
409
|
+
**From PyPI (no clone):**
|
|
410
|
+
|
|
411
|
+
```bash
|
|
412
|
+
pipx install devopsiq # or: pip install devopsiq
|
|
413
|
+
export OPENROUTER_API_KEY=sk-or-...
|
|
414
|
+
devopsiq --json "why is api-5d6f crash-looping?"
|
|
415
|
+
```
|
|
416
|
+
|
|
417
|
+
The console command is `devopsiq`; the importable package is
|
|
418
|
+
`agent` / `tools` / `main` (see INTEGRATION.md for library use).
|
|
419
|
+
|
|
420
|
+
**Prebuilt Docker image (bundles kubectl, helm, gh, trivy, git, curl,
|
|
421
|
+
docker CLI):**
|
|
422
|
+
|
|
423
|
+
```bash
|
|
424
|
+
docker run --rm \
|
|
425
|
+
-e OPENROUTER_API_KEY=sk-or-... \
|
|
426
|
+
-v "$HOME/.kube:/home/agent/.kube:ro" \
|
|
427
|
+
-v agent-records:/data \
|
|
428
|
+
ghcr.io/devopsabhii/devops-ai-agent --json "why is api-5d6f crash-looping?"
|
|
429
|
+
```
|
|
430
|
+
|
|
431
|
+
Images are multi-arch (amd64 + arm64), published on every `v*` tag
|
|
432
|
+
(`:latest` tracks the newest release). Add
|
|
433
|
+
`-v /var/run/docker.sock:/var/run/docker.sock` for the Docker/Compose tools.
|
|
434
|
+
|
|
435
|
+
**From source (development):**
|
|
436
|
+
|
|
437
|
+
Requires **Python 3.10+** (built-in `venv`). The tool layer itself uses only
|
|
438
|
+
the standard library; the OpenAI SDK and python-dotenv are the only Python
|
|
439
|
+
dependencies.
|
|
440
|
+
|
|
441
|
+
```bash
|
|
442
|
+
cd ~/devops-ai-agent
|
|
443
|
+
python3 -m venv .venv
|
|
444
|
+
.venv/bin/pip install --upgrade pip
|
|
445
|
+
.venv/bin/pip install -r requirements.txt
|
|
446
|
+
```
|
|
447
|
+
|
|
448
|
+
Recommended CLIs, by domain:
|
|
449
|
+
|
|
450
|
+
| Domain | CLIs needed |
|
|
451
|
+
| ------------ | ------------------------------------------------ |
|
|
452
|
+
| Kubernetes | `kubectl` (configured cluster context; metrics-server for `top`) |
|
|
453
|
+
| Linux system | `systemctl`, `journalctl`, `ss`, `ps` (systemd host) |
|
|
454
|
+
| Docker | `docker` (daemon running); `docker compose` for the compose tools |
|
|
455
|
+
| Terraform | `terraform` (initialized working directory) |
|
|
456
|
+
| Helm | `helm` (kubeconfig) |
|
|
457
|
+
| Argo CD | `argocd` (installed + `argocd login`) |
|
|
458
|
+
| Istio | `istioctl` (reachable mesh) |
|
|
459
|
+
| Security | `trivy` (first scan downloads the CVE DB) |
|
|
460
|
+
| git / GitHub | `git`; `gh` (authenticated, for PRs and Actions) |
|
|
461
|
+
| Cloud | `aws` / `gcloud` / `az` (only the ones you use) |
|
|
462
|
+
| Monitoring | `curl` + endpoint env vars (see below) |
|
|
463
|
+
| New Relic | `curl` + `NEW_RELIC_API_KEY` / `NEW_RELIC_ACCOUNT_ID` env vars |
|
|
464
|
+
| Ansible | `ansible-inventory`, `ansible-playbook` |
|
|
465
|
+
|
|
466
|
+
## 6. Environment setup
|
|
467
|
+
|
|
468
|
+
```bash
|
|
469
|
+
cp .env.example .env # then edit it
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
`.env` must contain your real key — get one at <https://openrouter.ai/keys>:
|
|
473
|
+
|
|
474
|
+
```
|
|
475
|
+
OPENROUTER_API_KEY=sk-or-...
|
|
476
|
+
```
|
|
477
|
+
|
|
478
|
+
Optional overrides (defaults shown):
|
|
479
|
+
|
|
480
|
+
```
|
|
481
|
+
OPENROUTER_BASE_URL=https://openrouter.ai/api/v1
|
|
482
|
+
OPENROUTER_MODEL=z-ai/glm-5.3
|
|
483
|
+
AGENT_STORE_DIR=~/.devops-ai-agent/investigations # (Phase 7) where records are saved
|
|
484
|
+
PROMETHEUS_URL=http://prometheus:9090 # (Phase 8) enables prom_query
|
|
485
|
+
LOKI_URL=http://loki:3100 # (Phase 8) enables loki_query
|
|
486
|
+
GRAFANA_URL=http://grafana:3000 # (Phase 8) enables grafana_health
|
|
487
|
+
NEW_RELIC_API_KEY=NRAK-... # (Phase 9) NerdGraph user key
|
|
488
|
+
NEW_RELIC_ACCOUNT_ID=1234567 # (Phase 9) numeric account id
|
|
489
|
+
```
|
|
490
|
+
|
|
491
|
+
`.env` is gitignored; the API key is never hard-coded in Python, printed, or
|
|
492
|
+
logged. If `OPENROUTER_API_KEY` is already set in your shell, the shell value
|
|
493
|
+
wins and `.env` is not consulted. Kubernetes tools use kubectl's own config
|
|
494
|
+
(`~/.kube/config` or `KUBECONFIG`); no agent-side config is needed.
|
|
495
|
+
|
|
496
|
+
## 7. How to run the agent
|
|
497
|
+
|
|
498
|
+
**REPL (interactive):**
|
|
499
|
+
|
|
500
|
+
```bash
|
|
501
|
+
cd ~/devops-ai-agent
|
|
502
|
+
.venv/bin/python main.py
|
|
503
|
+
```
|
|
504
|
+
|
|
505
|
+
**One-shot (cron / CI / scripts):** pass the problem as the first argument.
|
|
506
|
+
The agent runs once, prints its answer, and exits 0 on success, 1 on
|
|
507
|
+
setup/API errors — so it drops straight into a pipeline:
|
|
508
|
+
|
|
509
|
+
```bash
|
|
510
|
+
.venv/bin/python main.py "why is api-5d6f crash-looping?"
|
|
511
|
+
.venv/bin/python main.py --json "why is api-5d6f crash-looping?" # structured report
|
|
512
|
+
.venv/bin/python main.py --resume --json "any update?" # continue a prior run
|
|
513
|
+
.venv/bin/python main.py --store-dir /tmp/runs --out report.json "..." # pipeline paths
|
|
514
|
+
```
|
|
515
|
+
|
|
516
|
+
With `--json` the stdout is one JSON document (see `render_report_json` in
|
|
517
|
+
`agent/investigation.py`): `problem`, `status`, `hypotheses`, `evidence`,
|
|
518
|
+
and — once concluded — `conclusion` with `root_cause`, `remediation`,
|
|
519
|
+
`verification`, `confidence`. If the model never opened an investigation,
|
|
520
|
+
`--json`/`--out` fail with exit 1 rather than printing a malformed report.
|
|
521
|
+
`--out PATH` additionally writes that JSON to an exact path for pipelines
|
|
522
|
+
that want a known location; `--store-dir DIR` points persistence somewhere
|
|
523
|
+
else for the run; `--resume` continues the newest in-progress record from
|
|
524
|
+
the store before asking. Slash commands also work one-shot:
|
|
525
|
+
`python main.py /report`.
|
|
526
|
+
|
|
527
|
+
**Persistence (Phase 7):** every record mutation is auto-saved to
|
|
528
|
+
`~/.devops-ai-agent/investigations/` — one JSON file per investigation,
|
|
529
|
+
written atomically on every begin/record/conclude, so a crash mid-run
|
|
530
|
+
loses nothing. The REPL resumes the newest *in-progress* record at startup
|
|
531
|
+
(concluded records stay as history); `/investigations` lists everything on
|
|
532
|
+
disk; `/endinvestigation` clears memory but keeps the saved file. A broken
|
|
533
|
+
store (permissions, full disk) degrades to a warning in the tool result —
|
|
534
|
+
it never interrupts an investigation.
|
|
535
|
+
|
|
536
|
+
Example session:
|
|
537
|
+
|
|
538
|
+
```
|
|
539
|
+
DevOps AI Agent (Phase 9 — 58 read-only tools, persistent investigations)
|
|
540
|
+
Model: z-ai/glm-5.3
|
|
541
|
+
Backend: https://openrouter.ai/api/v1
|
|
542
|
+
Store: /home/you/.devops-ai-agent/investigations
|
|
543
|
+
Tools: ansible_inventory, ansible_playbook_tasks, aws_identity,
|
|
544
|
+
az_account, az_groups, docker_disk_usage, docker_images,
|
|
545
|
+
docker_inspect, docker_logs, docker_networks, docker_ps,
|
|
546
|
+
docker_stats, docker_volumes, gcloud_identity, gh_prs,
|
|
547
|
+
gh_run_view, gh_runs, gh_workflows, git_diff, git_log,
|
|
548
|
+
git_repo_status, grafana_health, investigation_begin,
|
|
549
|
+
investigation_conclude, investigation_record, k8s_contexts,
|
|
550
|
+
k8s_deployment_status, k8s_events, k8s_hpa, k8s_nodes,
|
|
551
|
+
k8s_pod_logs, k8s_pod_status, k8s_pods, k8s_pvc, k8s_services,
|
|
552
|
+
k8s_top_nodes, k8s_top_pods, loki_query, prom_query,
|
|
553
|
+
sys_open_ports, sys_service_logs, sys_service_status,
|
|
554
|
+
sys_top_processes, system_info, tf_plan, tf_show, tf_state_list
|
|
555
|
+
Commands: /investigate <problem>, /investigation, /investigations, /report, /endinvestigation
|
|
556
|
+
One-shot: python main.py [--json] [--resume] [--out report.json] [--store-dir DIR] "<problem>"
|
|
557
|
+
Type 'exit' to quit.
|
|
558
|
+
|
|
559
|
+
You: The checkout service container keeps exiting in Docker. Investigate.
|
|
560
|
+
Agent: (docker_ps → docker_inspect → docker_logs, tracks hypotheses, and
|
|
561
|
+
ends with the structured report: the crash command, exit code, and
|
|
562
|
+
the remediation recommendation)
|
|
563
|
+
|
|
564
|
+
You: /report
|
|
565
|
+
Agent: # Investigation report … (the canonical record, rendered from state)
|
|
566
|
+
|
|
567
|
+
You: exit
|
|
568
|
+
```
|
|
569
|
+
|
|
570
|
+
Type `exit` / `quit`, or press Ctrl-D / Ctrl-C to leave. Slash commands are
|
|
571
|
+
handled locally and never reach the model.
|
|
572
|
+
|
|
573
|
+
## 8. Current limitations (Phase 9)
|
|
574
|
+
|
|
575
|
+
- **Each domain needs its CLI installed and reachable.** Missing CLIs,
|
|
576
|
+
unauthenticated `gh`, a dead docker daemon, an uninitialized terraform
|
|
577
|
+
directory, or an unreachable cluster all return the exact error honestly —
|
|
578
|
+
nothing is invented, but a tool can't produce data without its backend.
|
|
579
|
+
- **terraform/git/gh/ansible tools are working-directory scoped.** They read
|
|
580
|
+
the directory the agent was launched from — no path arguments by design
|
|
581
|
+
(keeps traversal out). To investigate another repo/module, launch the
|
|
582
|
+
agent there.
|
|
583
|
+
- **Monitoring endpoints are environment-configured by design.** Unless the
|
|
584
|
+
operator sets `PROMETHEUS_URL` / `LOKI_URL` / `GRAFANA_URL`, those tools
|
|
585
|
+
fail with an honest message naming the variable; the agent never invents
|
|
586
|
+
a URL.
|
|
587
|
+
- **`kubectl top` needs metrics-server.** Clusters without it return
|
|
588
|
+
kubectl's exact error — honest, but no usage data.
|
|
589
|
+
- **Cloud tools are identity-level only.** `aws_identity`,
|
|
590
|
+
`gcloud_identity`, `az_account`, `az_groups` answer "which account am I
|
|
591
|
+
looking at" but do not sweep region-scoped resources (ec2 describe-*,
|
|
592
|
+
compute instances list, ...) yet.
|
|
593
|
+
- **Phase 9 tools need their CLIs.** `helm`, `argocd` (plus `argocd
|
|
594
|
+
login`), `istioctl` and `trivy` are not bundled; without them those
|
|
595
|
+
tools return the exact "not found" error. New Relic tools need
|
|
596
|
+
`NEW_RELIC_API_KEY` + `NEW_RELIC_ACCOUNT_ID` in the environment; without
|
|
597
|
+
them they name the missing variable — nothing is invented.
|
|
598
|
+
- **Trivy's first scan downloads the CVE database** and can take a couple
|
|
599
|
+
of minutes; subsequent scans are fast.
|
|
600
|
+
- **The model can only select from the registered tools.** It can never run
|
|
601
|
+
an arbitrary verb of any CLI, or a mutating one — by construction.
|
|
602
|
+
- **Raw CLI output to the model.** Python does not re-parse pod/docs/state;
|
|
603
|
+
the model interprets real output, truncated at 8,000 characters per result.
|
|
604
|
+
- **Conversation history is short-term only.** The investigation record
|
|
605
|
+
persists across CLI exits (Phase 7), but chat history still lives in the
|
|
606
|
+
process and is lost when the CLI exits.
|
|
607
|
+
- **Hypothesis tracking is model-driven.** The record is what the model
|
|
608
|
+
chose to record through the investigation tools; the live tracker returned
|
|
609
|
+
on every record call is designed to keep that complete, but it is still
|
|
610
|
+
the model's discipline.
|
|
611
|
+
- **No streaming, no retries/backoff** yet.
|
|
612
|
+
- **Read-only is enforced by construction today.** Mutating capabilities
|
|
613
|
+
will only be added behind an explicit human-approval gate, much later.
|
|
614
|
+
|
|
615
|
+
## 9. Planned future phases
|
|
616
|
+
|
|
617
|
+
- **Phase 2 — tool architecture. ✅ Done.** Tool contract, registry,
|
|
618
|
+
tool-use loop, first read-only tool, offline + live verification.
|
|
619
|
+
- **Phase 3 — Kubernetes tooling. ✅ Done.** Read-only `kubectl`-backed
|
|
620
|
+
tools (pod status/logs, deployment status), name validation, verb
|
|
621
|
+
allowlist, offline + live verification.
|
|
622
|
+
- **Phase 4 — investigation loop. ✅ Done.** The agent plans a multi-step
|
|
623
|
+
investigation, tracks hypotheses and evidence in a first-class in-memory
|
|
624
|
+
record, and produces a structured report (facts → observations →
|
|
625
|
+
hypotheses → conclusion → remediation → verification).
|
|
626
|
+
- **Phase 5 — multi-domain tooling. ✅ Done.** Linux system (systemd/
|
|
627
|
+
journal/ports/processes), Docker (ps/inspect/logs/stats/images),
|
|
628
|
+
Kubernetes depth (events/nodes/services), Terraform (show/state/plan),
|
|
629
|
+
git/GitHub (status/log/diff/PRs) — all read-only, offline + live verified.
|
|
630
|
+
- **Phase 6 — automation surface. ✅ Done.** One-shot CLI mode
|
|
631
|
+
(`python main.py [--json] "<problem>"`) with exit codes for cron/CI, and
|
|
632
|
+
a structured JSON export of the investigation report (`render_report_json`)
|
|
633
|
+
so pipelines can act on the verdict. Repository is git-tracked with a
|
|
634
|
+
clean `.gitignore` (secrets and the venv never ride along).
|
|
635
|
+
- **Phase 7 — persistence. ✅ Done.** The investigation record survives CLI
|
|
636
|
+
exits: auto-saved on every mutation to
|
|
637
|
+
`~/.devops-ai-agent/investigations/` (`--store-dir`/`AGENT_STORE_DIR` to
|
|
638
|
+
override), atomic in-place file updates, REPL auto-resume of the newest
|
|
639
|
+
in-progress record, `/investigations` listing, one-shot `--resume` /
|
|
640
|
+
`--out`, and lossless `to_dict`/`from_dict` round-tripping — all offline
|
|
641
|
+
tested.
|
|
642
|
+
- **Phase 8 — tool expansion. ✅ Done.** 21 more read-only tools (26 → 47):
|
|
643
|
+
Kubernetes depth (`k8s_pods`, `k8s_top_pods`, `k8s_top_nodes`, `k8s_hpa`,
|
|
644
|
+
`k8s_pvc`, `k8s_contexts`), Docker depth (`docker_networks`,
|
|
645
|
+
`docker_volumes`, `docker_disk_usage`), GitHub Actions (`gh_runs`,
|
|
646
|
+
`gh_run_view`, `gh_workflows`), cloud identity (`aws_identity`,
|
|
647
|
+
`gcloud_identity`, `az_account`, `az_groups`), monitoring/logging
|
|
648
|
+
(`prom_query`, `loki_query`, `grafana_health` — endpoints from env only),
|
|
649
|
+
and Ansible listing (`ansible_inventory`, `ansible_playbook_tasks`) —
|
|
650
|
+
offline tested, live-verified where the host has the CLI.
|
|
651
|
+
- **Phase 9 — trending-market tools. ✅ Done.** 11 more read-only tools
|
|
652
|
+
(47 → 58): New Relic (`newrelic_nrql`, `newrelic_alerts` — NerdGraph over
|
|
653
|
+
curl, credentials env-only, GraphQL variables so query text can't escape),
|
|
654
|
+
Trivy (`trivy_image_scan`), Helm (`helm_list`, `helm_status`,
|
|
655
|
+
`helm_history` — reads only), Argo CD (`argocd_apps`,
|
|
656
|
+
`argocd_app_status`), Istio (`istioctl_proxy_status`), Docker Compose
|
|
657
|
+
(`docker_compose_ls`, `docker_compose_ps` — live-verified on this host).
|
|
658
|
+
- **Later — region-scoped cloud resources** (ec2 describe-*, compute
|
|
659
|
+
instances list, ...) behind the same template pattern; more observability
|
|
660
|
+
depth (New Relic dashboards/entities, Prometheus range queries); streaming;
|
|
661
|
+
conversation-history persistence; and a human-approval gate before any
|
|
662
|
+
mutating action is ever allowed.
|