devopsiq 0.1.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,662 @@
1
+ Metadata-Version: 2.4
2
+ Name: devopsiq
3
+ Version: 0.1.0
4
+ Summary: A from-scratch, read-only AI agent that investigates real DevOps problems — Kubernetes, Docker, Helm, Argo CD, Istio, Terraform, GitHub Actions, cloud, New Relic and more.
5
+ Author: DevOpsAbhii
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/DevOpsAbhii/devops-ai-agent
8
+ Project-URL: Repository, https://github.com/DevOpsAbhii/devops-ai-agent
9
+ Project-URL: Issues, https://github.com/DevOpsAbhii/devops-ai-agent/issues
10
+ Keywords: devops,kubernetes,docker,helm,argocd,istio,terraform,github-actions,observability,newrelic,incident-response,root-cause-analysis,ai-agent,cli
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Environment :: Console
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Intended Audience :: System Administrators
15
+ Classifier: Operating System :: POSIX :: Linux
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3.10
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: System :: Systems Administration
22
+ Classifier: Topic :: Software Development :: Debuggers
23
+ Requires-Python: >=3.10
24
+ Description-Content-Type: text/markdown
25
+ License-File: LICENSE
26
+ Requires-Dist: openai>=1.40.0
27
+ Requires-Dist: python-dotenv>=1.0.1
28
+ Dynamic: license-file
29
+
30
+ # DevOps AI Agent
31
+
32
+ An AI agent that investigates real DevOps problems. The end goal: ask it
33
+ something like *"Why is my Kubernetes pod in CrashLoopBackOff?"* and have it
34
+ gather evidence, reason about the evidence, identify the likely root cause,
35
+ recommend remediation, and give verification steps.
36
+
37
+ **Phase 5** broadens the agent from
38
+ Kubernetes-only to the wider DevOps surface: Linux systemd services and
39
+ journals, Docker containers, Terraform state/plan, and git/GitHub pull
40
+ requests — every tool still read-only, still a fixed allowlisted argv
41
+ template. **Phase 6 adds the automation surface:** a one-shot CLI mode
42
+ (`python main.py "problem"` — exit-code driven, cron/CI-friendly) and a
43
+ structured JSON export of the investigation report, so pipelines can act on
44
+ the verdict instead of parsing markdown. **Phase 7 adds persistence:** the
45
+ investigation record is auto-saved on every change to
46
+ `~/.devops-ai-agent/investigations/` (overridable with `--store-dir` or
47
+ `AGENT_STORE_DIR`), survives CLI exits, and the REPL resumes the newest
48
+ in-progress record at startup. **Phase 8 broadens the tool surface to 47
49
+ read-only tools:** Kubernetes depth (pod listing, `kubectl top` resource
50
+ usage, HPAs, PVCs, contexts), Docker depth (networks, volumes, disk usage),
51
+ GitHub Actions via `gh` (runs, run jobs, workflows), cloud identity
52
+ (AWS/GCP/Azure), monitoring/logging (Prometheus, Loki, Grafana — endpoints
53
+ from environment config), and Ansible listing (inventory, playbook tasks).
54
+ **Phase 9 adds 11 trending-market tools (58 total):** New Relic (NRQL +
55
+ open alerts over NerdGraph — credentials from env), Trivy image
56
+ vulnerability scanning, Helm releases (list/status/history), Argo CD
57
+ (GitOps app sync/health), Istio mesh proxy status, and Docker Compose
58
+ (project and service listing). The repository is git-tracked.
59
+ Every phase still built from scratch — no LangChain, LangGraph,
60
+ AutoGen, CrewAI, or MCP.
61
+
62
+ > **Want to use this in your own project?** See [INTEGRATION.md](INTEGRATION.md)
63
+ > for the three integration levels: drive it as a CLI from cron/CI (exit
64
+ > codes + `--json`), embed it as a Python library, or extend it with your
65
+ > own read-only tools.
66
+
67
+ ---
68
+
69
+ ## 1. What this project is
70
+
71
+ The agent is built intentionally: we own the core architecture — how the
72
+ model is called, how conversation history flows, how tool selection +
73
+ execution + evidence feedback work — instead of depending on a framework
74
+ for it.
75
+
76
+ In Phase 5 the agent:
77
+ - holds a conversation with **GLM 5.3** through **OpenRouter**;
78
+ - has **58 real, read-only tools** across fifteen domains: host facts,
79
+ Kubernetes (12 tools), Linux system (4), Docker + Compose (10),
80
+ Terraform (3), Helm (3), Argo CD (2), Istio (1), security/Trivy (1),
81
+ git/GitHub + Actions (7), cloud identity (4), monitoring/logging (3),
82
+ New Relic (2), Ansible (2), plus the 3 investigation tools;
83
+ - runs a **tool-use loop**: when the model decides a question needs evidence,
84
+ it requests the tool, the application executes it locally, and the real
85
+ output is fed back to the model, which then answers from it;
86
+ - runs an **investigation loop**: for a reported problem it plans a multi-step
87
+ investigation — open a record with initial hypotheses, gather evidence with
88
+ the right domain tools, update each hypothesis's verdict as evidence
89
+ builds, conclude with root cause + remediation + verification, and end with
90
+ a structured report. The tracked state lives in the application, so
91
+ `/investigation` and `/report` show it directly.
92
+
93
+ It is **read-only by design**: every tool is a static, allowlisted command
94
+ template — the model can never pass arbitrary command text, can never select
95
+ a mutating verb (no delete/restart/edit/apply/scale/exec/run/rm/reset/
96
+ destroy anywhere), and no mutating capability exists. Object and unit names
97
+ are validated before reaching any CLI; flags are injected only by fixed
98
+ templates, never by model text. The investigation tools mutate only the
99
+ agent's in-memory record. Any future mutating capability will only ever
100
+ arrive behind an explicit human-approval gate.
101
+
102
+ ## 2. Current architecture
103
+
104
+ ```
105
+ You (terminal)
106
+ │ plain text
107
+ ▼
108
+ main.py ......................... CLI REPL, .env loading, error handling
109
+ │
110
+ ▼
111
+ agent/agent.py (DevOpsAgent) .... system prompt + conversation history
112
+ │
113
+ ▼ the tool-use loop, inside DevOpsAgent._complete():
114
+ │
115
+ │ attach tool schemas -> call model
116
+ │ │
117
+ │ ├─ model requests tool(s)? ── yes ─▶ tools/registry.execute_tool()
118
+ │ │ │ per call, in order
119
+ │ │ │ └─ allowlisted read-only
120
+ │ │ │ command runs for real
121
+ │ │ │ └─ result echoed back as
122
+ │ │ │ a "tool" message
123
+ │ │ ◀── loop calls the model again
124
+ │ └─ model answers in plain text? ── yes ─▶ done
125
+ │ (safety cap: max 10 tool-use turns, then abort)
126
+ ▼
127
+ OpenRouter (https://openrouter.ai/api/v1)
128
+ ▼
129
+ GLM 5.3 (z-ai/glm-5.3)
130
+ ```
131
+
132
+ Module map:
133
+
134
+ | Path | Responsibility |
135
+ | --------------------------- | ------------------------------------------------------------ |
136
+ | `main.py` | REPL loop, one-shot CLI, environment loading, `/investigate` commands, error messages |
137
+ | `agent/agent.py` | `DevOpsAgent` — client, history, `ask()`, `_complete()` loop |
138
+ | `agent/prompts.py` | The system prompt (versioned/tested separately) |
139
+ | `agent/investigation.py` | First-class investigation record: hypotheses, verdicts, evidence, report renderer (pure data) |
140
+ | `agent/store.py` | `InvestigationStore` — one JSON file per record, atomic writes, resume/list (Phase 7) |
141
+ | `tools/base.py` | Tool contract: `Tool`, `ToolError`, `read_command_output()` |
142
+ | `tools/registry.py` | `register/get_tools/execute_tool` — tools declared & executed here |
143
+ | `tools/preflight.py` | `system_info` — host facts (allowlisted read-only commands) |
144
+ | `tools/kubernetes.py` | 12 kubectl tools: pods, pod status/logs, deployments, events, nodes, top (cpu/mem), hpa, pvc, services, contexts |
145
+ | `tools/system.py` | systemd/journal/ss/ps tools — services, journals, ports, top processes |
146
+ | `tools/docker.py` | read-only docker tools — ps, inspect, logs, stats, images, networks, volumes, disk usage, compose ls/ps |
147
+ | `tools/terraform.py` | `tf_show` / `tf_state_list` / `tf_plan` (working directory) |
148
+ | `tools/helm.py` | `helm_list` / `helm_status` / `helm_history` — release reads only (Phase 9) |
149
+ | `tools/argocd.py` | `argocd_apps` / `argocd_app_status` — GitOps app reads (Phase 9) |
150
+ | `tools/istio.py` | `istioctl_proxy_status` — mesh sync view, no arguments (Phase 9) |
151
+ | `tools/trivy.py` | `trivy_image_scan` — vulnerability report for one image (Phase 9) |
152
+ | `tools/newrelic.py` | `newrelic_nrql` / `newrelic_alerts` — NerdGraph over curl, credentials from env only (Phase 9) |
153
+ | `tools/git_ci.py` | git status/log/diff + `gh_prs`, `gh_runs`, `gh_run_view`, `gh_workflows` (working directory) |
154
+ | `tools/cloud.py` | read-only cloud identity/listing: `aws_identity`, `gcloud_identity`, `az_account`, `az_groups` |
155
+ | `tools/monitoring.py` | `prom_query`, `loki_query`, `grafana_health` — endpoints from env config only |
156
+ | `tools/ansible.py` | `ansible_inventory`, `ansible_playbook_tasks` — listing modes only |
157
+ | `tools/investigation.py` | `investigation_begin` / `investigation_record` / `investigation_conclude` — meta-tools for the investigation record |
158
+ | `tests/test_phase2.py` | Offline suite: contract, safety, loop (stdlib `unittest`) |
159
+ | `tests/test_phase3.py` | Offline suite: k8s command lines + verb-allowlist (fake kubectl) |
160
+ | `tests/test_phase4.py` | Offline suite: investigation state, meta-tools, loop-driven investigation |
161
+ | `tests/test_phase5.py` | Offline suite: argv templates + name validation for all new domains (fake CLIs) |
162
+ | `tests/test_automation.py` | Offline suite: JSON reports + one-shot CLI (Phase 6) |
163
+ | `tests/test_phase7.py` | Offline suite: round-trip serialization, store files, auto-save, resume, CLI flags (Phase 7) |
164
+ | `tests/test_phase8.py` | Offline suite: argv templates + validation for the 21 Phase 8 tools (fake CLIs, env-based monitoring) |
165
+ | `tests/test_phase9.py` | Offline suite: New Relic env-credential + payload tests, trivy/helm/argocd/istio/compose argv templates (Phase 9) |
166
+
167
+ ### The tool-use loop
168
+
169
+ All model calls funnel through `DevOpsAgent._complete()`. Each turn:
170
+
171
+ 1. The full history is sent with every tool's schema attached (`tools=`).
172
+ 2. If the reply contains tool calls, the assistant tool-request turn is
173
+ echoed into history verbatim, and each call is executed locally via
174
+ `tools/registry.execute_tool(name, arguments)`.
175
+ 3. Every real result is appended as a sibling `{"role": "tool"}` message
176
+ pinned to its `tool_call_id`.
177
+ 4. The loop calls the model again — giving it the evidence to reflect on —
178
+ and repeats until it answers in plain text.
179
+ 5. `MAX_TOOL_ITERATIONS = 10` aborts runaway loops.
180
+
181
+ ### The investigation loop (Phase 4)
182
+
183
+ For a reported problem (a pod in CrashLoopBackOff, a broken rollout) the
184
+ same loop carries a *plan* as well as evidence:
185
+
186
+ 1. The model opens a record with `investigation_begin` (one-line problem +
187
+ 2–4 initial hypotheses, which become H1, H2, …).
188
+ 2. It plans what it needs, then calls the read-only tools one deliberate
189
+ step at a time.
190
+ 3. Every finding is recorded with `investigation_record`: evidence notes
191
+ (linked to a hypothesis when they bear on one) and verdicts
192
+ (supported / refuted / confirmed).
193
+ 4. Each record call returns the **live tracker** (`Problem`, status,
194
+ hypotheses with verdicts, evidence), so the model always knows where it
195
+ stands without inspecting all of history.
196
+ 5. When evidence is sufficient the model calls `investigation_conclude`
197
+ with root cause, remediation *recommendations* (nothing is ever
198
+ executed), verification steps, and confidence — then ends its answer
199
+ with the structured report.
200
+
201
+ The record lives in the **application**, not just the conversation: the
202
+ report is rendered deterministically from it (agent/investigation.py), and
203
+ the CLI exposes it directly:
204
+
205
+ | Command | What it does |
206
+ | ------------------------ | --------------------------------------------------- |
207
+ | `/investigate <problem>` | Open a formal investigation (same path the model uses) |
208
+ | `/investigation` | Show the live tracked state (hypotheses/evidence) |
209
+ | `/investigations` | List saved records on disk (newest first; `← active` marks the live one) |
210
+ | `/report` | Show the canonical report (once concluded) |
211
+ | `/endinvestigation` | Clear the record (memory only — the saved copy stays as history) |
212
+
213
+ The report the agent ends with and `/report` render are kept consistent by
214
+ construction: `tool` results confirm each record call and the tracker, and
215
+ `render_report()` is the single report format.
216
+
217
+ Example (live, against a real cluster):
218
+
219
+ ```
220
+ You: /investigate Pod crashloop-6f8c… is in CrashLoopBackOff in default ns.
221
+ (…) investigate.
222
+ Agent: (investigation_begin → k8s_pod_status → k8s_pod_logs →
223
+ k8s_deployment_status → verdicts on H1/H2 → investigation_conclude)
224
+ ends with: Facts → Hypotheses → Root cause → Remediation (recommended)
225
+ → Verification steps.
226
+
227
+ You: /report
228
+ Agent: # Investigation report … (canonical, rendered from the record)
229
+ ```
230
+
231
+ ### The tools (Phases 5, 8 and 9)
232
+
233
+ Every tool runs a fixed, allowlisted command — the model only picks
234
+ arguments (names, counts, namespaces), never command text.
235
+
236
+ | Domain | Tool | Backing command (read-only) |
237
+ | ---------------- | --------------------- | ---------------------------------------------------- |
238
+ | Host | `system_info` | `date`/`uname`/`uptime`/`df`/`free` |
239
+ | Kubernetes | `k8s_pod_status` | `kubectl get pod <pod> -n <ns> -o json` |
240
+ | | `k8s_pod_logs` | `kubectl logs <pod> -n <ns> --tail=<n>` |
241
+ | | `k8s_deployment_status` | `kubectl get deployment <dep> -n <ns> -o json` |
242
+ | | `k8s_events` | `kubectl get events -n <ns> --sort-by=.lastTimestamp -o wide [--field-selector involvedObject.name=<obj>]` |
243
+ | | `k8s_nodes` | `kubectl get nodes -o json` |
244
+ | | `k8s_services` | `kubectl get services -n <ns> -o json` |
245
+ | | `k8s_pods` | `kubectl get pods -n <ns> -o wide` |
246
+ | | `k8s_top_pods` | `kubectl top pods -n <ns> [--sort-by=cpu\|memory]` |
247
+ | | `k8s_top_nodes` | `kubectl top nodes` |
248
+ | | `k8s_hpa` | `kubectl get hpa [<name>] -n <ns> -o json` |
249
+ | | `k8s_pvc` | `kubectl get pvc [<name>] -n <ns> -o json` |
250
+ | | `k8s_contexts` | `kubectl config get-contexts` (listing only — never `use-context`) |
251
+ | Linux system | `sys_service_status` | `systemctl status <unit> --no-pager` |
252
+ | | `sys_service_logs` | `journalctl -u <unit> --no-pager -n <n>` |
253
+ | | `sys_open_ports` | `ss -tlnp` |
254
+ | | `sys_top_processes` | `ps aux --sort=-%cpu --no-headers` |
255
+ | Docker | `docker_ps` | `docker ps -a` |
256
+ | | `docker_inspect` | `docker inspect <name>` |
257
+ | | `docker_logs` | `docker logs --tail <n> <name>` |
258
+ | | `docker_stats` | `docker stats --no-stream` (the flag is pinned — see below) |
259
+ | | `docker_images` | `docker images` |
260
+ | | `docker_networks` | `docker network ls` |
261
+ | | `docker_volumes` | `docker volume ls` |
262
+ | | `docker_disk_usage` | `docker system df` |
263
+ | | `docker_compose_ls` | `docker compose ls` |
264
+ | | `docker_compose_ps` | `docker compose [-p <project>] ps -a` |
265
+ | Helm | `helm_list` | `helm list -n <ns>` / `helm list --all-namespaces` |
266
+ | | `helm_status` | `helm status <release> -n <ns>` |
267
+ | | `helm_history` | `helm history <release> -n <ns> --max <n>` |
268
+ | Argo CD | `argocd_apps` | `argocd app list --output json` |
269
+ | | `argocd_app_status` | `argocd app get <app>` |
270
+ | Istio | `istioctl_proxy_status` | `istioctl proxy-status` (no arguments) |
271
+ | Security | `trivy_image_scan` | `trivy image --scanners vuln --format table <image>` |
272
+ | New Relic | `newrelic_nrql` | `curl -H "API-Key: …" -d <json payload> https://api.newrelic.com/graphql` |
273
+ | | `newrelic_alerts` | same pinned curl + payload (no model input) |
274
+ | Terraform | `tf_show` | `terraform show -no-color` |
275
+ | | `tf_state_list` | `terraform state list` |
276
+ | | `tf_plan` | `terraform plan -no-color -input=false` |
277
+ | git / GitHub | `git_repo_status` | `git status --short --branch` |
278
+ | | `git_log` | `git log --oneline -n <n>` |
279
+ | | `git_diff` | `git diff --stat HEAD` |
280
+ | | `gh_prs` | `gh pr list --limit <n> --json number,title,state,...` |
281
+ | | `gh_runs` | `gh run list --limit <n> --json databaseId,displayTitle,status,...` |
282
+ | | `gh_run_view` | `gh run view <id> --json status,conclusion,jobs` (id digits-only) |
283
+ | | `gh_workflows` | `gh workflow list --limit <n> --json id,name,state` |
284
+ | Cloud identity | `aws_identity` | `aws sts get-caller-identity --output json` |
285
+ | | `gcloud_identity` | `gcloud config list --format=json` |
286
+ | | `az_account` | `az account show` |
287
+ | | `az_groups` | `az group list` |
288
+ | Monitoring | `prom_query` | `curl <PROMETHEUS_URL>/api/v1/query?query=<urlencoded PromQL>` |
289
+ | | `loki_query` | `curl <LOKI_URL>/loki/api/v1/query?query=<urlencoded LogQL>&limit=<n>` |
290
+ | | `grafana_health` | `curl <GRAFANA_URL>/api/health` |
291
+ | Ansible | `ansible_inventory` | `ansible-inventory [--inventory <source>] --list` |
292
+ | | `ansible_playbook_tasks` | `ansible-playbook --list-tasks --list-hosts <playbook>` |
293
+ | Investigation | `investigation_begin` / `investigation_record` / `investigation_conclude` | record only — memory, no external command |
294
+
295
+ Notes:
296
+
297
+ - Output is the **real** CLI output — Python does not re-parse it; the model
298
+ reads exactly what an engineer would see. Verbatim quoting is a hard
299
+ prompt rule.
300
+ - `docker stats` is **always** `--no-stream` (the pinned flag is what keeps
301
+ it from following forever and hanging the turn).
302
+ - `terraform plan` is a **dry run** — it computes the diff, mutates nothing;
303
+ `-input=false` keeps it from ever prompting. Requires an initialized
304
+ directory (`terraform init`); on a fresh dir the model reports terraform's
305
+ own error, honestly.
306
+ - Terraform and git/gh tools operate on the **current working directory** the
307
+ agent was launched from (no path parameters — that keeps traversal out).
308
+ - `gh_prs` needs the `gh` CLI authenticated and a GitHub remote; otherwise
309
+ the exact CLI error is returned.
310
+ - `kubectl top` needs **metrics-server** in the cluster; where it is absent,
311
+ kubectl's exact error is returned (that is honest behavior, not a bug).
312
+ - `k8s_contexts` is a listing of kubeconfig contexts only — the agent can
313
+ never run `config use-context` and change which cluster is in use.
314
+ - **Monitoring endpoints come from the environment, never the model:**
315
+ `prom_query`/`loki_query`/`grafana_health` read `PROMETHEUS_URL` /
316
+ `LOKI_URL` / `GRAFANA_URL`; unset or non-http(s) values are honest
317
+ ToolErrors naming the variable. The model supplies only the PromQL/LogQL
318
+ text, which is percent-encoded into the query string; the pinned curl argv
319
+ (`--proto =https,http`, GET only) blocks `file://` and everything but the
320
+ configured host.
321
+ - Cloud tools are identity/account level only and take no arguments at all;
322
+ Ansible tools are listing modes only (`--list`, `--list-tasks`) with
323
+ relative-path validation (no `/`, no leading `-`, no `..`).
324
+ - **Helm tools only read** (list/status/history — never install/upgrade/
325
+ rollback/uninstall); **Argo CD tools only list/get** (never sync/rollback/
326
+ delete); `istioctl_proxy_status` takes no arguments at all;
327
+ `trivy_image_scan` validates the image reference (no leading `-`, no
328
+ spaces) and warns its first run may take minutes (CVE DB download).
329
+ - **New Relic credentials are environment-configured, never model-chosen:**
330
+ `NEW_RELIC_API_KEY` (passed only as a curl header value) and
331
+ `NEW_RELIC_ACCOUNT_ID` (digits only). The NRQL text travels inside a
332
+ `json.dumps`-built payload as a GraphQL *variable*, so it can never
333
+ escape its string slot; only NerdGraph queries are ever sent, never
334
+ mutations.
335
+ - Log tails (`--tail`, `-n`, `--limit`) are bounded integers, validated
336
+ 1–500 (git log 1–100, gh 1–50, Loki limit 1–1000, helm history 1–50).
337
+ - Any missing CLI / unreachable target returns the exact error — nothing is
338
+ invented (verified live for kubectl, systemctl, journalctl, docker,
339
+ terraform, git).
340
+
341
+ ### How read-only is enforced (defense in depth)
342
+
343
+ - The tool schemas only allow picking names/counts/namespaces from validated
344
+ arguments — there is no way to pass command text to any CLI.
345
+ - Every tool is a fixed argv template. No `sh -c` anywhere, so nothing is
346
+ ever parsed by a shell; flags the model might abuse (e.g. `--no-stream`
347
+ for docker stats, `-input=false` for terraform plan) are hard-coded into
348
+ the template and cannot be removed or added.
349
+ - Only read-only verbs exist per domain — `get`/`logs`/`top` and
350
+ `config get-contexts` (kubectl), `status`/`-u`/`ss`/`ps` (system),
351
+ `ps`/`inspect`/`logs`/`stats`/`images`/`network ls`/`volume ls`/
352
+ `system df`/`compose ls`/`compose ps` (docker), `show`/`state list`/
353
+ `plan` (terraform), `list`/`status`/`history` (helm), `app list`/
354
+ `app get` (argocd), `proxy-status` (istioctl), `image --scanners vuln`
355
+ (trivy), `status`/`log`/`diff`/`pr list`/`run list`/`run view`/
356
+ `workflow list` (git/gh), `--list`/`--list-tasks` (ansible), GET-only
357
+ curl with a pinned argv (monitoring), query-only NerdGraph with a pinned
358
+ curl (New Relic). No delete, restart, edit, apply, scale, exec, run, rm,
359
+ pull, push, commit, reset, merge, use-context, playbook-run, install,
360
+ upgrade, rollback, uninstall, sync, destroy — by construction.
361
+ - Names are validated per domain before reaching any CLI: Kubernetes object
362
+ names (DNS style), systemd unit names (no `/`, no leading `-`), Docker
363
+ names (no `/`), GitHub run ids (digits only), Ansible sources (relative
364
+ paths, no `..`), and no path parameters at all for terraform/git/gh tools.
365
+ This blocks flag and path injection.
366
+ - The executor re-validates every argument. Never trust the model.
367
+ - `--request-timeout`/timeouts bound slow or hanging commands, and logs are
368
+ tail-bounded.
369
+ - Unrecognized tools/arguments return `Tool error: ...` to the model instead
370
+ of executing.
371
+ - Tool output is truncated at 8,000 characters per result.
372
+
373
+ ## 3. Why GLM 5.3
374
+
375
+ - GLM 5.3 (`z-ai/glm-5.3` on OpenRouter) is a capable, cost-effective
376
+ general model — with strong instruction-following and **function-calling**
377
+ support, which the tool loop has demonstrated live in every phase.
378
+ - Served by Zhipu AI through a single OpenRouter endpoint, so there is no
379
+ separate vendor API to manage.
380
+ - An investigation loop sends many tokens (system prompt, tool schemas, tool
381
+ output, history); GLM 5.3 keeps that cost sustainable.
382
+
383
+ ## 4. How OpenRouter fits
384
+
385
+ [OpenRouter](https://openrouter.ai) is a model gateway: one API key, one
386
+ OpenAI-compatible endpoint, access to many models. We call
387
+
388
+ ```
389
+ POST https://openrouter.ai/api/v1/chat/completions
390
+ ```
391
+
392
+ with model `z-ai/glm-5.3` and standard `tools` / `tool` messages. The
393
+ official `openai` Python SDK works as our client with two config lines:
394
+
395
+ ```python
396
+ OpenAI(api_key=..., base_url="https://openrouter.ai/api/v1")
397
+ ```
398
+
399
+ Swapping to another OpenRouter model later is a one-line change (an env var
400
+ today); moving to any other OpenAI-compatible provider changes only
401
+ `agent/agent.py` configuration.
402
+
403
+ ## 5. Installation
404
+
405
+ Three ways — pick one. Whatever you choose, the agent also needs the CLIs of
406
+ the domains you use (kubectl, docker, gh, … — each tool reports its exact
407
+ error if its CLI is missing, so a partial install is fine).
408
+
409
+ **From PyPI (no clone):**
410
+
411
+ ```bash
412
+ pipx install devopsiq # or: pip install devopsiq
413
+ export OPENROUTER_API_KEY=sk-or-...
414
+ devopsiq --json "why is api-5d6f crash-looping?"
415
+ ```
416
+
417
+ The console command is `devopsiq`; the importable package is
418
+ `agent` / `tools` / `main` (see INTEGRATION.md for library use).
419
+
420
+ **Prebuilt Docker image (bundles kubectl, helm, gh, trivy, git, curl,
421
+ docker CLI):**
422
+
423
+ ```bash
424
+ docker run --rm \
425
+ -e OPENROUTER_API_KEY=sk-or-... \
426
+ -v "$HOME/.kube:/home/agent/.kube:ro" \
427
+ -v agent-records:/data \
428
+ ghcr.io/devopsabhii/devops-ai-agent --json "why is api-5d6f crash-looping?"
429
+ ```
430
+
431
+ Images are multi-arch (amd64 + arm64), published on every `v*` tag
432
+ (`:latest` tracks the newest release). Add
433
+ `-v /var/run/docker.sock:/var/run/docker.sock` for the Docker/Compose tools.
434
+
435
+ **From source (development):**
436
+
437
+ Requires **Python 3.10+** (built-in `venv`). The tool layer itself uses only
438
+ the standard library; the OpenAI SDK and python-dotenv are the only Python
439
+ dependencies.
440
+
441
+ ```bash
442
+ cd ~/devops-ai-agent
443
+ python3 -m venv .venv
444
+ .venv/bin/pip install --upgrade pip
445
+ .venv/bin/pip install -r requirements.txt
446
+ ```
447
+
448
+ Recommended CLIs, by domain:
449
+
450
+ | Domain | CLIs needed |
451
+ | ------------ | ------------------------------------------------ |
452
+ | Kubernetes | `kubectl` (configured cluster context; metrics-server for `top`) |
453
+ | Linux system | `systemctl`, `journalctl`, `ss`, `ps` (systemd host) |
454
+ | Docker | `docker` (daemon running); `docker compose` for the compose tools |
455
+ | Terraform | `terraform` (initialized working directory) |
456
+ | Helm | `helm` (kubeconfig) |
457
+ | Argo CD | `argocd` (installed + `argocd login`) |
458
+ | Istio | `istioctl` (reachable mesh) |
459
+ | Security | `trivy` (first scan downloads the CVE DB) |
460
+ | git / GitHub | `git`; `gh` (authenticated, for PRs and Actions) |
461
+ | Cloud | `aws` / `gcloud` / `az` (only the ones you use) |
462
+ | Monitoring | `curl` + endpoint env vars (see below) |
463
+ | New Relic | `curl` + `NEW_RELIC_API_KEY` / `NEW_RELIC_ACCOUNT_ID` env vars |
464
+ | Ansible | `ansible-inventory`, `ansible-playbook` |
465
+
466
+ ## 6. Environment setup
467
+
468
+ ```bash
469
+ cp .env.example .env # then edit it
470
+ ```
471
+
472
+ `.env` must contain your real key — get one at <https://openrouter.ai/keys>:
473
+
474
+ ```
475
+ OPENROUTER_API_KEY=sk-or-...
476
+ ```
477
+
478
+ Optional overrides (defaults shown):
479
+
480
+ ```
481
+ OPENROUTER_BASE_URL=https://openrouter.ai/api/v1
482
+ OPENROUTER_MODEL=z-ai/glm-5.3
483
+ AGENT_STORE_DIR=~/.devops-ai-agent/investigations # (Phase 7) where records are saved
484
+ PROMETHEUS_URL=http://prometheus:9090 # (Phase 8) enables prom_query
485
+ LOKI_URL=http://loki:3100 # (Phase 8) enables loki_query
486
+ GRAFANA_URL=http://grafana:3000 # (Phase 8) enables grafana_health
487
+ NEW_RELIC_API_KEY=NRAK-... # (Phase 9) NerdGraph user key
488
+ NEW_RELIC_ACCOUNT_ID=1234567 # (Phase 9) numeric account id
489
+ ```
490
+
491
+ `.env` is gitignored; the API key is never hard-coded in Python, printed, or
492
+ logged. If `OPENROUTER_API_KEY` is already set in your shell, the shell value
493
+ wins and `.env` is not consulted. Kubernetes tools use kubectl's own config
494
+ (`~/.kube/config` or `KUBECONFIG`); no agent-side config is needed.
495
+
496
+ ## 7. How to run the agent
497
+
498
+ **REPL (interactive):**
499
+
500
+ ```bash
501
+ cd ~/devops-ai-agent
502
+ .venv/bin/python main.py
503
+ ```
504
+
505
+ **One-shot (cron / CI / scripts):** pass the problem as the first argument.
506
+ The agent runs once, prints its answer, and exits 0 on success, 1 on
507
+ setup/API errors — so it drops straight into a pipeline:
508
+
509
+ ```bash
510
+ .venv/bin/python main.py "why is api-5d6f crash-looping?"
511
+ .venv/bin/python main.py --json "why is api-5d6f crash-looping?" # structured report
512
+ .venv/bin/python main.py --resume --json "any update?" # continue a prior run
513
+ .venv/bin/python main.py --store-dir /tmp/runs --out report.json "..." # pipeline paths
514
+ ```
515
+
516
+ With `--json` the stdout is one JSON document (see `render_report_json` in
517
+ `agent/investigation.py`): `problem`, `status`, `hypotheses`, `evidence`,
518
+ and — once concluded — `conclusion` with `root_cause`, `remediation`,
519
+ `verification`, `confidence`. If the model never opened an investigation,
520
+ `--json`/`--out` fail with exit 1 rather than printing a malformed report.
521
+ `--out PATH` additionally writes that JSON to an exact path for pipelines
522
+ that want a known location; `--store-dir DIR` points persistence somewhere
523
+ else for the run; `--resume` continues the newest in-progress record from
524
+ the store before asking. Slash commands also work one-shot:
525
+ `python main.py /report`.
526
+
527
+ **Persistence (Phase 7):** every record mutation is auto-saved to
528
+ `~/.devops-ai-agent/investigations/` — one JSON file per investigation,
529
+ written atomically on every begin/record/conclude, so a crash mid-run
530
+ loses nothing. The REPL resumes the newest *in-progress* record at startup
531
+ (concluded records stay as history); `/investigations` lists everything on
532
+ disk; `/endinvestigation` clears memory but keeps the saved file. A broken
533
+ store (permissions, full disk) degrades to a warning in the tool result —
534
+ it never interrupts an investigation.
535
+
536
+ Example session:
537
+
538
+ ```
539
+ DevOps AI Agent (Phase 9 — 58 read-only tools, persistent investigations)
540
+ Model: z-ai/glm-5.3
541
+ Backend: https://openrouter.ai/api/v1
542
+ Store: /home/you/.devops-ai-agent/investigations
543
+ Tools: ansible_inventory, ansible_playbook_tasks, aws_identity,
544
+ az_account, az_groups, docker_disk_usage, docker_images,
545
+ docker_inspect, docker_logs, docker_networks, docker_ps,
546
+ docker_stats, docker_volumes, gcloud_identity, gh_prs,
547
+ gh_run_view, gh_runs, gh_workflows, git_diff, git_log,
548
+ git_repo_status, grafana_health, investigation_begin,
549
+ investigation_conclude, investigation_record, k8s_contexts,
550
+ k8s_deployment_status, k8s_events, k8s_hpa, k8s_nodes,
551
+ k8s_pod_logs, k8s_pod_status, k8s_pods, k8s_pvc, k8s_services,
552
+ k8s_top_nodes, k8s_top_pods, loki_query, prom_query,
553
+ sys_open_ports, sys_service_logs, sys_service_status,
554
+ sys_top_processes, system_info, tf_plan, tf_show, tf_state_list
555
+ Commands: /investigate <problem>, /investigation, /investigations, /report, /endinvestigation
556
+ One-shot: python main.py [--json] [--resume] [--out report.json] [--store-dir DIR] "<problem>"
557
+ Type 'exit' to quit.
558
+
559
+ You: The checkout service container keeps exiting in Docker. Investigate.
560
+ Agent: (docker_ps → docker_inspect → docker_logs, tracks hypotheses, and
561
+ ends with the structured report: the crash command, exit code, and
562
+ the remediation recommendation)
563
+
564
+ You: /report
565
+ Agent: # Investigation report … (the canonical record, rendered from state)
566
+
567
+ You: exit
568
+ ```
569
+
570
+ Type `exit` / `quit`, or press Ctrl-D / Ctrl-C to leave. Slash commands are
571
+ handled locally and never reach the model.
572
+
573
+ ## 8. Current limitations (Phase 9)
574
+
575
+ - **Each domain needs its CLI installed and reachable.** Missing CLIs,
576
+ unauthenticated `gh`, a dead docker daemon, an uninitialized terraform
577
+ directory, or an unreachable cluster all return the exact error honestly —
578
+ nothing is invented, but a tool can't produce data without its backend.
579
+ - **terraform/git/gh/ansible tools are working-directory scoped.** They read
580
+ the directory the agent was launched from — no path arguments by design
581
+ (keeps traversal out). To investigate another repo/module, launch the
582
+ agent there.
583
+ - **Monitoring endpoints are environment-configured by design.** Unless the
584
+ operator sets `PROMETHEUS_URL` / `LOKI_URL` / `GRAFANA_URL`, those tools
585
+ fail with an honest message naming the variable; the agent never invents
586
+ a URL.
587
+ - **`kubectl top` needs metrics-server.** Clusters without it return
588
+ kubectl's exact error — honest, but no usage data.
589
+ - **Cloud tools are identity-level only.** `aws_identity`,
590
+ `gcloud_identity`, `az_account`, `az_groups` answer "which account am I
591
+ looking at" but do not sweep region-scoped resources (ec2 describe-*,
592
+ compute instances list, ...) yet.
593
+ - **Phase 9 tools need their CLIs.** `helm`, `argocd` (plus `argocd
594
+ login`), `istioctl` and `trivy` are not bundled; without them those
595
+ tools return the exact "not found" error. New Relic tools need
596
+ `NEW_RELIC_API_KEY` + `NEW_RELIC_ACCOUNT_ID` in the environment; without
597
+ them they name the missing variable — nothing is invented.
598
+ - **Trivy's first scan downloads the CVE database** and can take a couple
599
+ of minutes; subsequent scans are fast.
600
+ - **The model can only select from the registered tools.** It can never run
601
+ an arbitrary verb of any CLI, or a mutating one — by construction.
602
+ - **Raw CLI output to the model.** Python does not re-parse pod/docs/state;
603
+ the model interprets real output, truncated at 8,000 characters per result.
604
+ - **Conversation history is short-term only.** The investigation record
605
+ persists across CLI exits (Phase 7), but chat history still lives in the
606
+ process and is lost when the CLI exits.
607
+ - **Hypothesis tracking is model-driven.** The record is what the model
608
+ chose to record through the investigation tools; the live tracker returned
609
+ on every record call is designed to keep that complete, but it is still
610
+ the model's discipline.
611
+ - **No streaming, no retries/backoff** yet.
612
+ - **Read-only is enforced by construction today.** Mutating capabilities
613
+ will only be added behind an explicit human-approval gate, much later.
614
+
615
+ ## 9. Planned future phases
616
+
617
+ - **Phase 2 — tool architecture. ✅ Done.** Tool contract, registry,
618
+ tool-use loop, first read-only tool, offline + live verification.
619
+ - **Phase 3 — Kubernetes tooling. ✅ Done.** Read-only `kubectl`-backed
620
+ tools (pod status/logs, deployment status), name validation, verb
621
+ allowlist, offline + live verification.
622
+ - **Phase 4 — investigation loop. ✅ Done.** The agent plans a multi-step
623
+ investigation, tracks hypotheses and evidence in a first-class in-memory
624
+ record, and produces a structured report (facts → observations →
625
+ hypotheses → conclusion → remediation → verification).
626
+ - **Phase 5 — multi-domain tooling. ✅ Done.** Linux system (systemd/
627
+ journal/ports/processes), Docker (ps/inspect/logs/stats/images),
628
+ Kubernetes depth (events/nodes/services), Terraform (show/state/plan),
629
+ git/GitHub (status/log/diff/PRs) — all read-only, offline + live verified.
630
+ - **Phase 6 — automation surface. ✅ Done.** One-shot CLI mode
631
+ (`python main.py [--json] "<problem>"`) with exit codes for cron/CI, and
632
+ a structured JSON export of the investigation report (`render_report_json`)
633
+ so pipelines can act on the verdict. Repository is git-tracked with a
634
+ clean `.gitignore` (secrets and the venv never ride along).
635
+ - **Phase 7 — persistence. ✅ Done.** The investigation record survives CLI
636
+ exits: auto-saved on every mutation to
637
+ `~/.devops-ai-agent/investigations/` (`--store-dir`/`AGENT_STORE_DIR` to
638
+ override), atomic in-place file updates, REPL auto-resume of the newest
639
+ in-progress record, `/investigations` listing, one-shot `--resume` /
640
+ `--out`, and lossless `to_dict`/`from_dict` round-tripping — all offline
641
+ tested.
642
+ - **Phase 8 — tool expansion. ✅ Done.** 21 more read-only tools (26 → 47):
643
+ Kubernetes depth (`k8s_pods`, `k8s_top_pods`, `k8s_top_nodes`, `k8s_hpa`,
644
+ `k8s_pvc`, `k8s_contexts`), Docker depth (`docker_networks`,
645
+ `docker_volumes`, `docker_disk_usage`), GitHub Actions (`gh_runs`,
646
+ `gh_run_view`, `gh_workflows`), cloud identity (`aws_identity`,
647
+ `gcloud_identity`, `az_account`, `az_groups`), monitoring/logging
648
+ (`prom_query`, `loki_query`, `grafana_health` — endpoints from env only),
649
+ and Ansible listing (`ansible_inventory`, `ansible_playbook_tasks`) —
650
+ offline tested, live-verified where the host has the CLI.
651
+ - **Phase 9 — trending-market tools. ✅ Done.** 11 more read-only tools
652
+ (47 → 58): New Relic (`newrelic_nrql`, `newrelic_alerts` — NerdGraph over
653
+ curl, credentials env-only, GraphQL variables so query text can't escape),
654
+ Trivy (`trivy_image_scan`), Helm (`helm_list`, `helm_status`,
655
+ `helm_history` — reads only), Argo CD (`argocd_apps`,
656
+ `argocd_app_status`), Istio (`istioctl_proxy_status`), Docker Compose
657
+ (`docker_compose_ls`, `docker_compose_ps` — live-verified on this host).
658
+ - **Later — region-scoped cloud resources** (ec2 describe-*, compute
659
+ instances list, ...) behind the same template pattern; more observability
660
+ depth (New Relic dashboards/entities, Prometheus range queries); streaming;
661
+ conversation-history persistence; and a human-approval gate before any
662
+ mutating action is ever allowed.