opencode-skills-collection 4.0.68 → 4.0.69
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +266 -1
- package/bundled-skills/access-review/SKILL.md +394 -0
- package/bundled-skills/access-review/references/details.md +121 -0
- package/bundled-skills/agent-evals/SKILL.md +420 -0
- package/bundled-skills/agent-observability/SKILL.md +346 -0
- package/bundled-skills/agent-observability/references/details.md +786 -0
- package/bundled-skills/ai-agent-security/SKILL.md +393 -0
- package/bundled-skills/ai-agent-security/references/details.md +912 -0
- package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
- package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
- package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
- package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
- package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
- package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
- package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
- package/bundled-skills/alerting-oncall/SKILL.md +458 -0
- package/bundled-skills/alerting-oncall/references/details.md +84 -0
- package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
- package/bundled-skills/argocd-gitops/SKILL.md +469 -0
- package/bundled-skills/arm-templates/SKILL.md +438 -0
- package/bundled-skills/arm-templates/references/details.md +64 -0
- package/bundled-skills/asset-inventory/SKILL.md +412 -0
- package/bundled-skills/asset-inventory/references/details.md +127 -0
- package/bundled-skills/audit-logging/SKILL.md +476 -0
- package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
- package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
- package/bundled-skills/aws-ec2/SKILL.md +426 -0
- package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
- package/bundled-skills/aws-iam/SKILL.md +463 -0
- package/bundled-skills/aws-lambda/SKILL.md +428 -0
- package/bundled-skills/aws-rds/SKILL.md +380 -0
- package/bundled-skills/aws-s3/SKILL.md +434 -0
- package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
- package/bundled-skills/aws-vpc/SKILL.md +436 -0
- package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
- package/bundled-skills/azure-aks/SKILL.md +423 -0
- package/bundled-skills/azure-devops/SKILL.md +457 -0
- package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
- package/bundled-skills/azure-keyvault/SKILL.md +455 -0
- package/bundled-skills/azure-keyvault/references/details.md +83 -0
- package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
- package/bundled-skills/azure-networking/SKILL.md +448 -0
- package/bundled-skills/azure-networking/references/details.md +135 -0
- package/bundled-skills/azure-sql/SKILL.md +413 -0
- package/bundled-skills/azure-sql/references/details.md +113 -0
- package/bundled-skills/azure-vms/SKILL.md +402 -0
- package/bundled-skills/azure-vms/references/details.md +134 -0
- package/bundled-skills/backup-recovery/SKILL.md +388 -0
- package/bundled-skills/bb-methodology/SKILL.md +451 -0
- package/bundled-skills/bb-methodology/references/details.md +120 -0
- package/bundled-skills/block-storage/SKILL.md +371 -0
- package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
- package/bundled-skills/blue-green-deploy/references/details.md +90 -0
- package/bundled-skills/bug-bounty/SKILL.md +447 -0
- package/bundled-skills/bug-bounty/references/details.md +1316 -0
- package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
- package/bundled-skills/business-continuity/SKILL.md +463 -0
- package/bundled-skills/career-ops/SKILL.md +186 -0
- package/bundled-skills/cdn-setup/SKILL.md +374 -0
- package/bundled-skills/change-management/SKILL.md +438 -0
- package/bundled-skills/change-management/references/details.md +105 -0
- package/bundled-skills/circleci/SKILL.md +475 -0
- package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
- package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
- package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
- package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
- package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
- package/bundled-skills/cloudformation/SKILL.md +461 -0
- package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
- package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
- package/bundled-skills/container-hardening/SKILL.md +126 -0
- package/bundled-skills/container-registries/SKILL.md +435 -0
- package/bundled-skills/container-scanning/SKILL.md +416 -0
- package/bundled-skills/convex-backend/SKILL.md +338 -0
- package/bundled-skills/dast-scanning/SKILL.md +437 -0
- package/bundled-skills/database-backups/SKILL.md +425 -0
- package/bundled-skills/datadog/SKILL.md +487 -0
- package/bundled-skills/dependency-scanning/SKILL.md +457 -0
- package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
- package/bundled-skills/disaster-recovery/SKILL.md +374 -0
- package/bundled-skills/disaster-recovery/references/details.md +219 -0
- package/bundled-skills/dns-management/SKILL.md +375 -0
- package/bundled-skills/docker-compose/SKILL.md +482 -0
- package/bundled-skills/docker-management/SKILL.md +426 -0
- package/bundled-skills/ebpf-observability/SKILL.md +436 -0
- package/bundled-skills/ebpf-observability/references/details.md +542 -0
- package/bundled-skills/elk-stack/SKILL.md +487 -0
- package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
- package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
- package/bundled-skills/feature-flags/SKILL.md +426 -0
- package/bundled-skills/feature-flags/references/details.md +86 -0
- package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
- package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
- package/bundled-skills/firewall-config/SKILL.md +479 -0
- package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
- package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
- package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
- package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
- package/bundled-skills/gcp-compute/SKILL.md +319 -0
- package/bundled-skills/gcp-gke/SKILL.md +307 -0
- package/bundled-skills/gcp-networking/SKILL.md +293 -0
- package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
- package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
- package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
- package/bundled-skills/gdpr-compliance/references/details.md +145 -0
- package/bundled-skills/geo-audit/SKILL.md +368 -0
- package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
- package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
- package/bundled-skills/geo-citability/SKILL.md +350 -0
- package/bundled-skills/geo-compare/SKILL.md +340 -0
- package/bundled-skills/geo-content/SKILL.md +383 -0
- package/bundled-skills/geo-crawlers/SKILL.md +408 -0
- package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
- package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
- package/bundled-skills/geo-proposal/SKILL.md +378 -0
- package/bundled-skills/geo-prospect/SKILL.md +225 -0
- package/bundled-skills/geo-report/SKILL.md +436 -0
- package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
- package/bundled-skills/geo-schema/SKILL.md +408 -0
- package/bundled-skills/geo-technical/SKILL.md +78 -0
- package/bundled-skills/geo-technical/references/details.md +543 -0
- package/bundled-skills/git-workflow/SKILL.md +460 -0
- package/bundled-skills/github-actions/SKILL.md +368 -0
- package/bundled-skills/gitlab-ci/SKILL.md +340 -0
- package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
- package/bundled-skills/gpu-server-management/SKILL.md +236 -0
- package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
- package/bundled-skills/helm-charts/SKILL.md +469 -0
- package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
- package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
- package/bundled-skills/hunt-ato/SKILL.md +184 -0
- package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
- package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
- package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
- package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
- package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
- package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
- package/bundled-skills/hunt-cicd/SKILL.md +311 -0
- package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
- package/bundled-skills/hunt-cors/SKILL.md +335 -0
- package/bundled-skills/hunt-dom/SKILL.md +323 -0
- package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
- package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
- package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
- package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
- package/bundled-skills/hunt-grpc/SKILL.md +317 -0
- package/bundled-skills/hunt-host-header/SKILL.md +309 -0
- package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
- package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
- package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
- package/bundled-skills/hunt-idor/SKILL.md +434 -0
- package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
- package/bundled-skills/hunt-k8s/SKILL.md +337 -0
- package/bundled-skills/hunt-laravel/SKILL.md +255 -0
- package/bundled-skills/hunt-ldap/SKILL.md +351 -0
- package/bundled-skills/hunt-lfi/SKILL.md +311 -0
- package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
- package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
- package/bundled-skills/hunt-misc/SKILL.md +378 -0
- package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
- package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
- package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
- package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
- package/bundled-skills/hunt-oauth/SKILL.md +459 -0
- package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
- package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
- package/bundled-skills/hunt-race-condition/references/details.md +159 -0
- package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
- package/bundled-skills/hunt-rce/SKILL.md +444 -0
- package/bundled-skills/hunt-rce/references/details.md +110 -0
- package/bundled-skills/hunt-saml/SKILL.md +156 -0
- package/bundled-skills/hunt-session/SKILL.md +342 -0
- package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
- package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
- package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
- package/bundled-skills/hunt-springboot/SKILL.md +285 -0
- package/bundled-skills/hunt-sqli/SKILL.md +466 -0
- package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
- package/bundled-skills/hunt-ssrf/references/details.md +179 -0
- package/bundled-skills/hunt-ssti/SKILL.md +163 -0
- package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
- package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
- package/bundled-skills/hunt-xxe/SKILL.md +466 -0
- package/bundled-skills/i-have-adhd/SKILL.md +170 -0
- package/bundled-skills/identity-access-management/SKILL.md +382 -0
- package/bundled-skills/identity-access-management/references/details.md +524 -0
- package/bundled-skills/incident-management/SKILL.md +484 -0
- package/bundled-skills/incident-response/SKILL.md +448 -0
- package/bundled-skills/incident-response/references/details.md +113 -0
- package/bundled-skills/interview-me/SKILL.md +248 -0
- package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
- package/bundled-skills/jenkins/SKILL.md +462 -0
- package/bundled-skills/jev-use/SKILL.md +158 -0
- package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
- package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
- package/bundled-skills/kubernetes-ops/references/details.md +108 -0
- package/bundled-skills/kustomize/SKILL.md +478 -0
- package/bundled-skills/linux-administration/SKILL.md +367 -0
- package/bundled-skills/linux-hardening/SKILL.md +154 -0
- package/bundled-skills/llm-app-security/SKILL.md +389 -0
- package/bundled-skills/llm-app-security/references/details.md +674 -0
- package/bundled-skills/llm-caching/SKILL.md +334 -0
- package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
- package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
- package/bundled-skills/llm-gateway/SKILL.md +282 -0
- package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
- package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
- package/bundled-skills/load-balancing/SKILL.md +403 -0
- package/bundled-skills/loki-logging/SKILL.md +479 -0
- package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
- package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
- package/bundled-skills/mcp-server-security/SKILL.md +356 -0
- package/bundled-skills/mcp-server-security/references/details.md +745 -0
- package/bundled-skills/mdm-device-management/SKILL.md +404 -0
- package/bundled-skills/mdm-device-management/references/details.md +410 -0
- package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
- package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
- package/bundled-skills/model-registry-governance/SKILL.md +452 -0
- package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
- package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
- package/bundled-skills/mongodb/SKILL.md +436 -0
- package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
- package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
- package/bundled-skills/mysql/SKILL.md +390 -0
- package/bundled-skills/new-relic/SKILL.md +472 -0
- package/bundled-skills/nfs-storage/SKILL.md +356 -0
- package/bundled-skills/object-storage/SKILL.md +378 -0
- package/bundled-skills/offensive-osint/SKILL.md +443 -0
- package/bundled-skills/okta-attack/SKILL.md +436 -0
- package/bundled-skills/ollama-stack/SKILL.md +379 -0
- package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
- package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
- package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
- package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
- package/bundled-skills/openshift/SKILL.md +485 -0
- package/bundled-skills/opentelemetry/SKILL.md +438 -0
- package/bundled-skills/opentelemetry/references/details.md +78 -0
- package/bundled-skills/opentofu-migration/SKILL.md +349 -0
- package/bundled-skills/osint-methodology/SKILL.md +460 -0
- package/bundled-skills/osint-methodology/references/details.md +1350 -0
- package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
- package/bundled-skills/penetration-testing/SKILL.md +152 -0
- package/bundled-skills/performance-tuning/SKILL.md +381 -0
- package/bundled-skills/planetscale/SKILL.md +297 -0
- package/bundled-skills/platform-engineering/SKILL.md +348 -0
- package/bundled-skills/platform-engineering/references/details.md +944 -0
- package/bundled-skills/podman/SKILL.md +405 -0
- package/bundled-skills/policy-as-code/SKILL.md +434 -0
- package/bundled-skills/policy-as-code/references/details.md +204 -0
- package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
- package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
- package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
- package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
- package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
- package/bundled-skills/rag-observability-evals/references/details.md +92 -0
- package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
- package/bundled-skills/redis/SKILL.md +421 -0
- package/bundled-skills/redteam-report-template/SKILL.md +370 -0
- package/bundled-skills/report-writing/SKILL.md +426 -0
- package/bundled-skills/report-writing/references/details.md +187 -0
- package/bundled-skills/reverse-proxy/SKILL.md +420 -0
- package/bundled-skills/runbook-creation/SKILL.md +438 -0
- package/bundled-skills/runbook-creation/references/details.md +71 -0
- package/bundled-skills/saas-security-posture/SKILL.md +415 -0
- package/bundled-skills/sast-scanning/SKILL.md +444 -0
- package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
- package/bundled-skills/security-arsenal/SKILL.md +446 -0
- package/bundled-skills/security-arsenal/references/details.md +540 -0
- package/bundled-skills/security-automation/SKILL.md +146 -0
- package/bundled-skills/semantic-versioning/SKILL.md +434 -0
- package/bundled-skills/semantic-versioning/references/details.md +83 -0
- package/bundled-skills/service-mesh/SKILL.md +422 -0
- package/bundled-skills/soc2-compliance/SKILL.md +409 -0
- package/bundled-skills/sops-encryption/SKILL.md +124 -0
- package/bundled-skills/sre-dashboards/SKILL.md +143 -0
- package/bundled-skills/ssh-configuration/SKILL.md +324 -0
- package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
- package/bundled-skills/ssl-tls-management/references/details.md +99 -0
- package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
- package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
- package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
- package/bundled-skills/systemd-services/SKILL.md +379 -0
- package/bundled-skills/terraform-aws/SKILL.md +125 -0
- package/bundled-skills/terraform-azure/SKILL.md +415 -0
- package/bundled-skills/terraform-azure/references/details.md +231 -0
- package/bundled-skills/terraform-gcp/SKILL.md +369 -0
- package/bundled-skills/threat-modeling/SKILL.md +487 -0
- package/bundled-skills/user-management/SKILL.md +383 -0
- package/bundled-skills/using-agent-skills/SKILL.md +220 -0
- package/bundled-skills/vector-database-ops/SKILL.md +300 -0
- package/bundled-skills/vendor-management/SKILL.md +439 -0
- package/bundled-skills/vendor-management/references/details.md +109 -0
- package/bundled-skills/vercel-deployments/SKILL.md +296 -0
- package/bundled-skills/vllm-server/SKILL.md +236 -0
- package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
- package/bundled-skills/vpn-setup/SKILL.md +452 -0
- package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
- package/bundled-skills/waf-setup/SKILL.md +354 -0
- package/bundled-skills/waf-setup/references/details.md +211 -0
- package/bundled-skills/web2-recon/SKILL.md +440 -0
- package/bundled-skills/web2-recon/references/details.md +319 -0
- package/bundled-skills/web3-audit/SKILL.md +445 -0
- package/bundled-skills/web3-audit/references/details.md +224 -0
- package/bundled-skills/windows-hardening/SKILL.md +454 -0
- package/bundled-skills/windows-hardening/references/details.md +204 -0
- package/bundled-skills/windows-server/SKILL.md +318 -0
- package/bundled-skills/zero-trust/SKILL.md +461 -0
- package/package.json +1 -1
- package/skills_index.json +6943 -323
|
@@ -0,0 +1,346 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: agent-observability
|
|
3
|
+
description: Instrument AI agents with tracing, token metrics, latency, and cost visibility.
|
|
4
|
+
Use for reliability and debugging.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
|
|
14
|
+
CI runners) and authorized access to the target environment. Docs-only; helper scripts
|
|
15
|
+
and templates not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '2.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# Agent Observability
|
|
22
|
+
|
|
23
|
+
Monitor AI agent behavior with logs, traces, metrics, and cost telemetry. This skill covers the full observability stack for LLM-powered applications: from raw Prometheus counters to Grafana dashboards, OpenTelemetry tracing, structured logging, cost tracking, SLO definition, and PII redaction.
|
|
24
|
+
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
## Core Metrics
|
|
28
|
+
|
|
29
|
+
Define these metrics at the application layer. All examples use the Prometheus client library naming conventions.
|
|
30
|
+
|
|
31
|
+
### Latency
|
|
32
|
+
|
|
33
|
+
```python
|
|
34
|
+
from prometheus_client import Histogram
|
|
35
|
+
|
|
36
|
+
# Total end-to-end latency for a full agent turn (user prompt -> final response)
|
|
37
|
+
AGENT_LATENCY = Histogram(
|
|
38
|
+
"agent_request_duration_seconds",
|
|
39
|
+
"End-to-end latency of an agent request",
|
|
40
|
+
labelnames=["agent_name", "model", "status"],
|
|
41
|
+
buckets=(0.25, 0.5, 1, 2, 5, 10, 30, 60, 120),
|
|
42
|
+
)
|
|
43
|
+
|
|
44
|
+
# Latency of a single LLM API call (one completion request)
|
|
45
|
+
LLM_CALL_LATENCY = Histogram(
|
|
46
|
+
"llm_call_duration_seconds",
|
|
47
|
+
"Latency of an individual LLM API call",
|
|
48
|
+
labelnames=["model", "provider", "stream"],
|
|
49
|
+
buckets=(0.1, 0.25, 0.5, 1, 2, 5, 10, 30),
|
|
50
|
+
)
|
|
51
|
+
|
|
52
|
+
# Latency of tool/function calls executed by the agent
|
|
53
|
+
TOOL_CALL_LATENCY = Histogram(
|
|
54
|
+
"agent_tool_call_duration_seconds",
|
|
55
|
+
"Latency of a tool call executed by the agent",
|
|
56
|
+
labelnames=["tool_name", "agent_name", "status"],
|
|
57
|
+
buckets=(0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10),
|
|
58
|
+
)
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
### Token Usage
|
|
62
|
+
|
|
63
|
+
```python
|
|
64
|
+
from prometheus_client import Counter, Histogram
|
|
65
|
+
|
|
66
|
+
PROMPT_TOKENS = Counter(
|
|
67
|
+
"llm_prompt_tokens_total",
|
|
68
|
+
"Total prompt tokens sent to the model",
|
|
69
|
+
labelnames=["model", "agent_name"],
|
|
70
|
+
)
|
|
71
|
+
|
|
72
|
+
COMPLETION_TOKENS = Counter(
|
|
73
|
+
"llm_completion_tokens_total",
|
|
74
|
+
"Total completion tokens received from the model",
|
|
75
|
+
labelnames=["model", "agent_name"],
|
|
76
|
+
)
|
|
77
|
+
|
|
78
|
+
CACHED_TOKENS = Counter(
|
|
79
|
+
"llm_cached_tokens_total",
|
|
80
|
+
"Prompt tokens served from KV-cache (provider-reported)",
|
|
81
|
+
labelnames=["model", "agent_name"],
|
|
82
|
+
)
|
|
83
|
+
|
|
84
|
+
TOKENS_PER_REQUEST = Histogram(
|
|
85
|
+
"llm_tokens_per_request",
|
|
86
|
+
"Total tokens (prompt + completion) per request",
|
|
87
|
+
labelnames=["model", "agent_name"],
|
|
88
|
+
buckets=(100, 500, 1000, 2000, 4000, 8000, 16000, 32000, 64000, 128000),
|
|
89
|
+
)
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
### Cost
|
|
93
|
+
|
|
94
|
+
```python
|
|
95
|
+
from prometheus_client import Counter
|
|
96
|
+
|
|
97
|
+
LLM_COST = Counter(
|
|
98
|
+
"llm_cost_dollars_total",
|
|
99
|
+
"Estimated cost in USD for LLM usage",
|
|
100
|
+
labelnames=["model", "agent_name", "cost_type"], # cost_type: prompt | completion
|
|
101
|
+
)
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
### Tool Calls
|
|
105
|
+
|
|
106
|
+
```python
|
|
107
|
+
from prometheus_client import Counter
|
|
108
|
+
|
|
109
|
+
TOOL_CALLS_TOTAL = Counter(
|
|
110
|
+
"agent_tool_calls_total",
|
|
111
|
+
"Total tool calls made by agents",
|
|
112
|
+
labelnames=["tool_name", "agent_name", "status"], # status: success | error | timeout
|
|
113
|
+
)
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### Errors and Retries
|
|
117
|
+
|
|
118
|
+
```python
|
|
119
|
+
from prometheus_client import Counter, Gauge
|
|
120
|
+
|
|
121
|
+
LLM_ERRORS = Counter(
|
|
122
|
+
"llm_errors_total",
|
|
123
|
+
"Errors returned by the LLM provider",
|
|
124
|
+
labelnames=["model", "provider", "error_type"], # error_type: rate_limit | timeout | 5xx | auth
|
|
125
|
+
)
|
|
126
|
+
|
|
127
|
+
LLM_RETRIES = Counter(
|
|
128
|
+
"llm_retries_total",
|
|
129
|
+
"Retried LLM API calls",
|
|
130
|
+
labelnames=["model", "provider", "retry_reason"],
|
|
131
|
+
)
|
|
132
|
+
|
|
133
|
+
AGENT_ACTIVE_REQUESTS = Gauge(
|
|
134
|
+
"agent_active_requests",
|
|
135
|
+
"Number of agent requests currently in flight",
|
|
136
|
+
labelnames=["agent_name"],
|
|
137
|
+
)
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
---
|
|
141
|
+
|
|
142
|
+
## OpenTelemetry Integration
|
|
143
|
+
|
|
144
|
+
Use the OpenTelemetry Python SDK to create traces that capture every step of an agent turn: the top-level request, each LLM call, each tool execution, and retrieval operations.
|
|
145
|
+
|
|
146
|
+
### Setup
|
|
147
|
+
|
|
148
|
+
```python
|
|
149
|
+
# otel_setup.py
|
|
150
|
+
from opentelemetry import trace
|
|
151
|
+
from opentelemetry.sdk.trace import TracerProvider
|
|
152
|
+
from opentelemetry.sdk.trace.export import BatchSpanProcessor
|
|
153
|
+
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
|
|
154
|
+
from opentelemetry.sdk.resources import Resource
|
|
155
|
+
|
|
156
|
+
def init_tracing(service_name: str, otlp_endpoint: str = "http://localhost:4317"):
|
|
157
|
+
resource = Resource.create({
|
|
158
|
+
"service.name": service_name,
|
|
159
|
+
"service.version": "1.0.0",
|
|
160
|
+
"deployment.environment": "production",
|
|
161
|
+
})
|
|
162
|
+
provider = TracerProvider(resource=resource)
|
|
163
|
+
exporter = OTLPSpanExporter(endpoint=otlp_endpoint, insecure=True)
|
|
164
|
+
provider.add_span_processor(BatchSpanProcessor(exporter))
|
|
165
|
+
trace.set_tracer_provider(provider)
|
|
166
|
+
return trace.get_tracer(service_name)
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
### Tracing LLM Calls
|
|
170
|
+
|
|
171
|
+
```python
|
|
172
|
+
# llm_tracing.py
|
|
173
|
+
import time
|
|
174
|
+
from opentelemetry import trace
|
|
175
|
+
from opentelemetry.trace import StatusCode
|
|
176
|
+
|
|
177
|
+
tracer = trace.get_tracer("agent.llm")
|
|
178
|
+
|
|
179
|
+
def traced_llm_call(client, messages, model="gpt-4o", **kwargs):
|
|
180
|
+
"""Wrap an LLM completion call with a full OpenTelemetry span."""
|
|
181
|
+
with tracer.start_as_current_span("llm.chat_completion") as span:
|
|
182
|
+
span.set_attribute("llm.model", model)
|
|
183
|
+
span.set_attribute("llm.provider", "openai")
|
|
184
|
+
span.set_attribute("llm.message_count", len(messages))
|
|
185
|
+
span.set_attribute("llm.temperature", kwargs.get("temperature", 1.0))
|
|
186
|
+
span.set_attribute("llm.max_tokens", kwargs.get("max_tokens", 0))
|
|
187
|
+
|
|
188
|
+
start = time.perf_counter()
|
|
189
|
+
try:
|
|
190
|
+
response = client.chat.completions.create(
|
|
191
|
+
model=model, messages=messages, **kwargs
|
|
192
|
+
)
|
|
193
|
+
elapsed = time.perf_counter() - start
|
|
194
|
+
|
|
195
|
+
usage = response.usage
|
|
196
|
+
span.set_attribute("llm.prompt_tokens", usage.prompt_tokens)
|
|
197
|
+
span.set_attribute("llm.completion_tokens", usage.completion_tokens)
|
|
198
|
+
span.set_attribute("llm.total_tokens", usage.total_tokens)
|
|
199
|
+
span.set_attribute("llm.duration_seconds", elapsed)
|
|
200
|
+
span.set_attribute("llm.finish_reason", response.choices[0].finish_reason)
|
|
201
|
+
span.set_status(StatusCode.OK)
|
|
202
|
+
|
|
203
|
+
# Update Prometheus counters
|
|
204
|
+
PROMPT_TOKENS.labels(model=model, agent_name="default").inc(usage.prompt_tokens)
|
|
205
|
+
COMPLETION_TOKENS.labels(model=model, agent_name="default").inc(usage.completion_tokens)
|
|
206
|
+
LLM_CALL_LATENCY.labels(model=model, provider="openai", stream="false").observe(elapsed)
|
|
207
|
+
|
|
208
|
+
return response
|
|
209
|
+
|
|
210
|
+
except Exception as exc:
|
|
211
|
+
elapsed = time.perf_counter() - start
|
|
212
|
+
span.set_status(StatusCode.ERROR, str(exc))
|
|
213
|
+
span.record_exception(exc)
|
|
214
|
+
LLM_ERRORS.labels(model=model, provider="openai", error_type=type(exc).__name__).inc()
|
|
215
|
+
raise
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
### Tracing Tool Execution
|
|
219
|
+
|
|
220
|
+
```python
|
|
221
|
+
# tool_tracing.py
|
|
222
|
+
import functools
|
|
223
|
+
from opentelemetry import trace
|
|
224
|
+
from opentelemetry.trace import StatusCode
|
|
225
|
+
|
|
226
|
+
tracer = trace.get_tracer("agent.tools")
|
|
227
|
+
|
|
228
|
+
def traced_tool(tool_name: str):
|
|
229
|
+
"""Decorator that wraps a tool function with an OTel span and Prometheus metrics."""
|
|
230
|
+
def decorator(func):
|
|
231
|
+
@functools.wraps(func)
|
|
232
|
+
def wrapper(*args, **kwargs):
|
|
233
|
+
with tracer.start_as_current_span(f"tool.{tool_name}") as span:
|
|
234
|
+
span.set_attribute("tool.name", tool_name)
|
|
235
|
+
span.set_attribute("tool.args_count", len(args) + len(kwargs))
|
|
236
|
+
|
|
237
|
+
import time
|
|
238
|
+
start = time.perf_counter()
|
|
239
|
+
try:
|
|
240
|
+
result = func(*args, **kwargs)
|
|
241
|
+
elapsed = time.perf_counter() - start
|
|
242
|
+
span.set_attribute("tool.duration_seconds", elapsed)
|
|
243
|
+
span.set_status(StatusCode.OK)
|
|
244
|
+
TOOL_CALLS_TOTAL.labels(
|
|
245
|
+
tool_name=tool_name, agent_name="default", status="success"
|
|
246
|
+
).inc()
|
|
247
|
+
TOOL_CALL_LATENCY.labels(
|
|
248
|
+
tool_name=tool_name, agent_name="default", status="success"
|
|
249
|
+
).observe(elapsed)
|
|
250
|
+
return result
|
|
251
|
+
except Exception as exc:
|
|
252
|
+
elapsed = time.perf_counter() - start
|
|
253
|
+
span.set_status(StatusCode.ERROR, str(exc))
|
|
254
|
+
span.record_exception(exc)
|
|
255
|
+
TOOL_CALLS_TOTAL.labels(
|
|
256
|
+
tool_name=tool_name, agent_name="default", status="error"
|
|
257
|
+
).inc()
|
|
258
|
+
TOOL_CALL_LATENCY.labels(
|
|
259
|
+
tool_name=tool_name, agent_name="default", status="error"
|
|
260
|
+
).observe(elapsed)
|
|
261
|
+
raise
|
|
262
|
+
return wrapper
|
|
263
|
+
return decorator
|
|
264
|
+
|
|
265
|
+
# Usage
|
|
266
|
+
@traced_tool("web_search")
|
|
267
|
+
def web_search(query: str) -> str:
|
|
268
|
+
# ... tool implementation ...
|
|
269
|
+
pass
|
|
270
|
+
|
|
271
|
+
@traced_tool("sql_query")
|
|
272
|
+
def sql_query(statement: str) -> list:
|
|
273
|
+
# ... tool implementation ...
|
|
274
|
+
pass
|
|
275
|
+
```
|
|
276
|
+
|
|
277
|
+
### Propagating Trace Context Across Services
|
|
278
|
+
|
|
279
|
+
```python
|
|
280
|
+
# context_propagation.py
|
|
281
|
+
from opentelemetry import context
|
|
282
|
+
from opentelemetry.propagate import inject, extract
|
|
283
|
+
import httpx
|
|
284
|
+
|
|
285
|
+
def call_downstream_service(url: str, payload: dict) -> dict:
|
|
286
|
+
"""Propagate the current trace context to a downstream HTTP service."""
|
|
287
|
+
headers = {}
|
|
288
|
+
inject(headers) # injects traceparent + tracestate headers
|
|
289
|
+
response = httpx.post(url, json=payload, headers=headers)
|
|
290
|
+
response.raise_for_status()
|
|
291
|
+
return response.json()
|
|
292
|
+
|
|
293
|
+
def extract_context_from_request(request_headers: dict):
|
|
294
|
+
"""Extract trace context from incoming request headers (for the receiving service)."""
|
|
295
|
+
ctx = extract(request_headers)
|
|
296
|
+
token = context.attach(ctx)
|
|
297
|
+
return token # call context.detach(token) when done
|
|
298
|
+
```
|
|
299
|
+
|
|
300
|
+
---
|
|
301
|
+
|
|
302
|
+
|
|
303
|
+
## Contents
|
|
304
|
+
|
|
305
|
+
- [Structured Logging](references/details.md)
|
|
306
|
+
- [Grafana Dashboards](references/details.md)
|
|
307
|
+
- [Cost Tracking](references/details.md)
|
|
308
|
+
- [Langfuse / Helicone Integration](references/details.md)
|
|
309
|
+
- [SLO Definition](references/details.md)
|
|
310
|
+
- [Debugging Workflows](references/details.md)
|
|
311
|
+
- [PII Redaction in Traces](references/details.md)
|
|
312
|
+
- [Best Practices](references/details.md)
|
|
313
|
+
- [Related Skills](references/details.md)
|
|
314
|
+
|
|
315
|
+
## When to Use
|
|
316
|
+
|
|
317
|
+
Apply this skill whenever you operate:
|
|
318
|
+
|
|
319
|
+
- **Autonomous AI agents** that make multi-step tool calls (e.g., coding agents, support agents, data-pipeline agents).
|
|
320
|
+
- **LLM-backed APIs** serving chat completions, summarisation, or classification behind a REST or gRPC gateway.
|
|
321
|
+
- **RAG pipelines** where a retriever fetches context from a vector store before prompting a model.
|
|
322
|
+
- **Multi-agent orchestrations** (crew-style or graph-based) where several agents collaborate on a single task.
|
|
323
|
+
- **Batch inference jobs** that process thousands of prompts against a model endpoint.
|
|
324
|
+
|
|
325
|
+
Key signals that you need this skill:
|
|
326
|
+
|
|
327
|
+
1. You cannot answer "what is p95 latency for agent responses this week?"
|
|
328
|
+
2. You have no per-request cost attribution.
|
|
329
|
+
3. Debugging a bad agent response requires grepping raw application logs.
|
|
330
|
+
4. You have no alerting on token-usage spikes or elevated error rates.
|
|
331
|
+
|
|
332
|
+
---
|
|
333
|
+
|
|
334
|
+
## Limitations
|
|
335
|
+
|
|
336
|
+
- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
|
|
337
|
+
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
|
|
338
|
+
|
|
339
|
+
### Example
|
|
340
|
+
|
|
341
|
+
```bash
|
|
342
|
+
git status && git diff --stat
|
|
343
|
+
kubectl diff -f manifest.yaml
|
|
344
|
+
```
|
|
345
|
+
|
|
346
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|