opencode-skills-collection 4.0.67 → 4.0.69
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +281 -1
- package/bundled-skills/access-review/SKILL.md +394 -0
- package/bundled-skills/access-review/references/details.md +121 -0
- package/bundled-skills/agent-evals/SKILL.md +420 -0
- package/bundled-skills/agent-observability/SKILL.md +346 -0
- package/bundled-skills/agent-observability/references/details.md +786 -0
- package/bundled-skills/ai-agent-security/SKILL.md +393 -0
- package/bundled-skills/ai-agent-security/references/details.md +912 -0
- package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
- package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
- package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
- package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
- package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
- package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
- package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
- package/bundled-skills/alerting-oncall/SKILL.md +458 -0
- package/bundled-skills/alerting-oncall/references/details.md +84 -0
- package/bundled-skills/anti-slop-design/SKILL.md +393 -0
- package/bundled-skills/antigravity-maintainer-batch-release/SKILL.md +1 -0
- package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
- package/bundled-skills/argocd-gitops/SKILL.md +469 -0
- package/bundled-skills/arm-templates/SKILL.md +438 -0
- package/bundled-skills/arm-templates/references/details.md +64 -0
- package/bundled-skills/artifact-yylo/SKILL.md +122 -0
- package/bundled-skills/asset-inventory/SKILL.md +412 -0
- package/bundled-skills/asset-inventory/references/details.md +127 -0
- package/bundled-skills/audit-logging/SKILL.md +476 -0
- package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
- package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
- package/bundled-skills/aws-ec2/SKILL.md +426 -0
- package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
- package/bundled-skills/aws-iam/SKILL.md +463 -0
- package/bundled-skills/aws-lambda/SKILL.md +428 -0
- package/bundled-skills/aws-rds/SKILL.md +380 -0
- package/bundled-skills/aws-s3/SKILL.md +434 -0
- package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
- package/bundled-skills/aws-vpc/SKILL.md +436 -0
- package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
- package/bundled-skills/azure-aks/SKILL.md +423 -0
- package/bundled-skills/azure-devops/SKILL.md +457 -0
- package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
- package/bundled-skills/azure-keyvault/SKILL.md +455 -0
- package/bundled-skills/azure-keyvault/references/details.md +83 -0
- package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
- package/bundled-skills/azure-networking/SKILL.md +448 -0
- package/bundled-skills/azure-networking/references/details.md +135 -0
- package/bundled-skills/azure-sql/SKILL.md +413 -0
- package/bundled-skills/azure-sql/references/details.md +113 -0
- package/bundled-skills/azure-vms/SKILL.md +402 -0
- package/bundled-skills/azure-vms/references/details.md +134 -0
- package/bundled-skills/backup-recovery/SKILL.md +388 -0
- package/bundled-skills/bb-methodology/SKILL.md +451 -0
- package/bundled-skills/bb-methodology/references/details.md +120 -0
- package/bundled-skills/block-storage/SKILL.md +371 -0
- package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
- package/bundled-skills/blue-green-deploy/references/details.md +90 -0
- package/bundled-skills/bug-bounty/SKILL.md +447 -0
- package/bundled-skills/bug-bounty/references/details.md +1316 -0
- package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
- package/bundled-skills/business-continuity/SKILL.md +463 -0
- package/bundled-skills/career-ops/SKILL.md +186 -0
- package/bundled-skills/cdn-setup/SKILL.md +374 -0
- package/bundled-skills/change-management/SKILL.md +438 -0
- package/bundled-skills/change-management/references/details.md +105 -0
- package/bundled-skills/circleci/SKILL.md +475 -0
- package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
- package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
- package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
- package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
- package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
- package/bundled-skills/cloudformation/SKILL.md +461 -0
- package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
- package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
- package/bundled-skills/container-hardening/SKILL.md +126 -0
- package/bundled-skills/container-registries/SKILL.md +435 -0
- package/bundled-skills/container-scanning/SKILL.md +416 -0
- package/bundled-skills/convex-backend/SKILL.md +338 -0
- package/bundled-skills/dast-scanning/SKILL.md +437 -0
- package/bundled-skills/database-backups/SKILL.md +425 -0
- package/bundled-skills/datadog/SKILL.md +487 -0
- package/bundled-skills/dependency-scanning/SKILL.md +457 -0
- package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
- package/bundled-skills/disaster-recovery/SKILL.md +374 -0
- package/bundled-skills/disaster-recovery/references/details.md +219 -0
- package/bundled-skills/dns-management/SKILL.md +375 -0
- package/bundled-skills/docker-compose/SKILL.md +482 -0
- package/bundled-skills/docker-management/SKILL.md +426 -0
- package/bundled-skills/ebpf-observability/SKILL.md +436 -0
- package/bundled-skills/ebpf-observability/references/details.md +542 -0
- package/bundled-skills/elk-stack/SKILL.md +487 -0
- package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
- package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
- package/bundled-skills/feature-flags/SKILL.md +426 -0
- package/bundled-skills/feature-flags/references/details.md +86 -0
- package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
- package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
- package/bundled-skills/firewall-config/SKILL.md +479 -0
- package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
- package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
- package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
- package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
- package/bundled-skills/gcp-compute/SKILL.md +319 -0
- package/bundled-skills/gcp-gke/SKILL.md +307 -0
- package/bundled-skills/gcp-networking/SKILL.md +293 -0
- package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
- package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
- package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
- package/bundled-skills/gdpr-compliance/references/details.md +145 -0
- package/bundled-skills/geo-audit/SKILL.md +368 -0
- package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
- package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
- package/bundled-skills/geo-citability/SKILL.md +350 -0
- package/bundled-skills/geo-compare/SKILL.md +340 -0
- package/bundled-skills/geo-content/SKILL.md +383 -0
- package/bundled-skills/geo-crawlers/SKILL.md +408 -0
- package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
- package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
- package/bundled-skills/geo-proposal/SKILL.md +378 -0
- package/bundled-skills/geo-prospect/SKILL.md +225 -0
- package/bundled-skills/geo-report/SKILL.md +436 -0
- package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
- package/bundled-skills/geo-schema/SKILL.md +408 -0
- package/bundled-skills/geo-technical/SKILL.md +78 -0
- package/bundled-skills/geo-technical/references/details.md +543 -0
- package/bundled-skills/git-workflow/SKILL.md +460 -0
- package/bundled-skills/github-actions/SKILL.md +368 -0
- package/bundled-skills/gitlab-ci/SKILL.md +340 -0
- package/bundled-skills/google-no-code/SKILL.md +136 -0
- package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
- package/bundled-skills/gpu-server-management/SKILL.md +236 -0
- package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
- package/bundled-skills/helm-charts/SKILL.md +469 -0
- package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
- package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
- package/bundled-skills/hunt-ato/SKILL.md +184 -0
- package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
- package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
- package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
- package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
- package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
- package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
- package/bundled-skills/hunt-cicd/SKILL.md +311 -0
- package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
- package/bundled-skills/hunt-cors/SKILL.md +335 -0
- package/bundled-skills/hunt-dom/SKILL.md +323 -0
- package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
- package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
- package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
- package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
- package/bundled-skills/hunt-grpc/SKILL.md +317 -0
- package/bundled-skills/hunt-host-header/SKILL.md +309 -0
- package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
- package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
- package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
- package/bundled-skills/hunt-idor/SKILL.md +434 -0
- package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
- package/bundled-skills/hunt-k8s/SKILL.md +337 -0
- package/bundled-skills/hunt-laravel/SKILL.md +255 -0
- package/bundled-skills/hunt-ldap/SKILL.md +351 -0
- package/bundled-skills/hunt-lfi/SKILL.md +311 -0
- package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
- package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
- package/bundled-skills/hunt-misc/SKILL.md +378 -0
- package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
- package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
- package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
- package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
- package/bundled-skills/hunt-oauth/SKILL.md +459 -0
- package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
- package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
- package/bundled-skills/hunt-race-condition/references/details.md +159 -0
- package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
- package/bundled-skills/hunt-rce/SKILL.md +444 -0
- package/bundled-skills/hunt-rce/references/details.md +110 -0
- package/bundled-skills/hunt-saml/SKILL.md +156 -0
- package/bundled-skills/hunt-session/SKILL.md +342 -0
- package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
- package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
- package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
- package/bundled-skills/hunt-springboot/SKILL.md +285 -0
- package/bundled-skills/hunt-sqli/SKILL.md +466 -0
- package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
- package/bundled-skills/hunt-ssrf/references/details.md +179 -0
- package/bundled-skills/hunt-ssti/SKILL.md +163 -0
- package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
- package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
- package/bundled-skills/hunt-xxe/SKILL.md +466 -0
- package/bundled-skills/i-have-adhd/SKILL.md +170 -0
- package/bundled-skills/idea-evaluator/SKILL.md +75 -0
- package/bundled-skills/idea-evaluator/idea-evaluator-con/SKILL.md +64 -0
- package/bundled-skills/idea-evaluator/idea-evaluator-pro/SKILL.md +64 -0
- package/bundled-skills/identity-access-management/SKILL.md +382 -0
- package/bundled-skills/identity-access-management/references/details.md +524 -0
- package/bundled-skills/incident-management/SKILL.md +484 -0
- package/bundled-skills/incident-response/SKILL.md +448 -0
- package/bundled-skills/incident-response/references/details.md +113 -0
- package/bundled-skills/interview-me/SKILL.md +248 -0
- package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
- package/bundled-skills/jenkins/SKILL.md +462 -0
- package/bundled-skills/jev-use/SKILL.md +158 -0
- package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
- package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
- package/bundled-skills/kubernetes-ops/references/details.md +108 -0
- package/bundled-skills/kustomize/SKILL.md +478 -0
- package/bundled-skills/ledger-tasks-yylo/SKILL.md +219 -0
- package/bundled-skills/linux-administration/SKILL.md +367 -0
- package/bundled-skills/linux-hardening/SKILL.md +154 -0
- package/bundled-skills/llm-app-security/SKILL.md +389 -0
- package/bundled-skills/llm-app-security/references/details.md +674 -0
- package/bundled-skills/llm-caching/SKILL.md +334 -0
- package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
- package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
- package/bundled-skills/llm-gateway/SKILL.md +282 -0
- package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
- package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
- package/bundled-skills/load-balancing/SKILL.md +403 -0
- package/bundled-skills/loki-logging/SKILL.md +479 -0
- package/bundled-skills/loki-mode/examples/todo-app-generated/backend/package-lock.json +4 -4
- package/bundled-skills/loki-mode/examples/todo-app-generated/backend/package.json +1 -1
- package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
- package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
- package/bundled-skills/mcp-server-security/SKILL.md +356 -0
- package/bundled-skills/mcp-server-security/references/details.md +745 -0
- package/bundled-skills/mdm-device-management/SKILL.md +404 -0
- package/bundled-skills/mdm-device-management/references/details.md +410 -0
- package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
- package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
- package/bundled-skills/model-registry-governance/SKILL.md +452 -0
- package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
- package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
- package/bundled-skills/mongodb/SKILL.md +436 -0
- package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
- package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
- package/bundled-skills/mysql/SKILL.md +390 -0
- package/bundled-skills/new-relic/SKILL.md +472 -0
- package/bundled-skills/nfs-storage/SKILL.md +356 -0
- package/bundled-skills/object-storage/SKILL.md +378 -0
- package/bundled-skills/offensive-osint/SKILL.md +443 -0
- package/bundled-skills/okta-attack/SKILL.md +436 -0
- package/bundled-skills/ollama-stack/SKILL.md +379 -0
- package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
- package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
- package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
- package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
- package/bundled-skills/openshift/SKILL.md +485 -0
- package/bundled-skills/opentelemetry/SKILL.md +438 -0
- package/bundled-skills/opentelemetry/references/details.md +78 -0
- package/bundled-skills/opentofu-migration/SKILL.md +349 -0
- package/bundled-skills/osint-methodology/SKILL.md +460 -0
- package/bundled-skills/osint-methodology/references/details.md +1350 -0
- package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
- package/bundled-skills/penetration-testing/SKILL.md +152 -0
- package/bundled-skills/performance-tuning/SKILL.md +381 -0
- package/bundled-skills/plan-ledger-tasks-yylo/SKILL.md +52 -0
- package/bundled-skills/planetscale/SKILL.md +297 -0
- package/bundled-skills/platform-engineering/SKILL.md +348 -0
- package/bundled-skills/platform-engineering/references/details.md +944 -0
- package/bundled-skills/podman/SKILL.md +405 -0
- package/bundled-skills/policy-as-code/SKILL.md +434 -0
- package/bundled-skills/policy-as-code/references/details.md +204 -0
- package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
- package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
- package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
- package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
- package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
- package/bundled-skills/rag-observability-evals/references/details.md +92 -0
- package/bundled-skills/ralph-loop-yylo/SKILL.md +55 -0
- package/bundled-skills/ralph-loop-yylo/references/first_check.md +18 -0
- package/bundled-skills/ralph-loop-yylo/references/implement.md +60 -0
- package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
- package/bundled-skills/redis/SKILL.md +421 -0
- package/bundled-skills/redteam-report-template/SKILL.md +370 -0
- package/bundled-skills/report-writing/SKILL.md +426 -0
- package/bundled-skills/report-writing/references/details.md +187 -0
- package/bundled-skills/resumable-implementation-contracts/SKILL.md +254 -0
- package/bundled-skills/reverse-proxy/SKILL.md +420 -0
- package/bundled-skills/runbook-creation/SKILL.md +438 -0
- package/bundled-skills/runbook-creation/references/details.md +71 -0
- package/bundled-skills/saas-security-posture/SKILL.md +415 -0
- package/bundled-skills/sast-scanning/SKILL.md +444 -0
- package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
- package/bundled-skills/security-arsenal/SKILL.md +446 -0
- package/bundled-skills/security-arsenal/references/details.md +540 -0
- package/bundled-skills/security-automation/SKILL.md +146 -0
- package/bundled-skills/semantic-versioning/SKILL.md +434 -0
- package/bundled-skills/semantic-versioning/references/details.md +83 -0
- package/bundled-skills/service-mesh/SKILL.md +422 -0
- package/bundled-skills/soc2-compliance/SKILL.md +409 -0
- package/bundled-skills/sops-encryption/SKILL.md +124 -0
- package/bundled-skills/sre-dashboards/SKILL.md +143 -0
- package/bundled-skills/ssh-configuration/SKILL.md +324 -0
- package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
- package/bundled-skills/ssl-tls-management/references/details.md +99 -0
- package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
- package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
- package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
- package/bundled-skills/systemd-services/SKILL.md +379 -0
- package/bundled-skills/terraform-aws/SKILL.md +125 -0
- package/bundled-skills/terraform-azure/SKILL.md +415 -0
- package/bundled-skills/terraform-azure/references/details.md +231 -0
- package/bundled-skills/terraform-gcp/SKILL.md +369 -0
- package/bundled-skills/threat-modeling/SKILL.md +487 -0
- package/bundled-skills/understand-project-yylo/SKILL.md +62 -0
- package/bundled-skills/user-management/SKILL.md +383 -0
- package/bundled-skills/using-agent-skills/SKILL.md +220 -0
- package/bundled-skills/vector-database-ops/SKILL.md +300 -0
- package/bundled-skills/vendor-management/SKILL.md +439 -0
- package/bundled-skills/vendor-management/references/details.md +109 -0
- package/bundled-skills/vercel-deployments/SKILL.md +296 -0
- package/bundled-skills/vllm-server/SKILL.md +236 -0
- package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
- package/bundled-skills/vpn-setup/SKILL.md +452 -0
- package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
- package/bundled-skills/waf-setup/SKILL.md +354 -0
- package/bundled-skills/waf-setup/references/details.md +211 -0
- package/bundled-skills/weather-model-data-fetching/SKILL.md +277 -0
- package/bundled-skills/weather-observation-fetching/SKILL.md +246 -0
- package/bundled-skills/web2-recon/SKILL.md +440 -0
- package/bundled-skills/web2-recon/references/details.md +319 -0
- package/bundled-skills/web3-audit/SKILL.md +445 -0
- package/bundled-skills/web3-audit/references/details.md +224 -0
- package/bundled-skills/wiki-yylo/SKILL.md +114 -0
- package/bundled-skills/windows-hardening/SKILL.md +454 -0
- package/bundled-skills/windows-hardening/references/details.md +204 -0
- package/bundled-skills/windows-server/SKILL.md +318 -0
- package/bundled-skills/workflow-yylo/SKILL.md +107 -0
- package/bundled-skills/zero-trust/SKILL.md +461 -0
- package/package.json +1 -1
- package/skills_index.json +6924 -0
|
@@ -0,0 +1,472 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: llmops-platform-engineering
|
|
3
|
+
description: Build production LLMOps platforms with CI/CD, model promotion workflows,
|
|
4
|
+
evaluation gates, rollback, and governance across cloud and self-hosted inference.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
|
|
14
|
+
CI runners) and authorized access to the target environment. Docs-only; helper scripts
|
|
15
|
+
and templates not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '1.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# LLMOps Platform Engineering
|
|
22
|
+
|
|
23
|
+
Design and operate an internal LLM platform that supports rapid experimentation without compromising reliability, cost, or compliance.
|
|
24
|
+
|
|
25
|
+
## When to Use This Skill
|
|
26
|
+
|
|
27
|
+
- Building an internal platform for teams to deploy and manage LLM-powered features
|
|
28
|
+
- Designing CI/CD pipelines that include model evaluation gates
|
|
29
|
+
- Setting up A/B testing infrastructure for model versions
|
|
30
|
+
- Creating Kubernetes-based model serving infrastructure
|
|
31
|
+
- Establishing governance workflows for model promotion
|
|
32
|
+
|
|
33
|
+
## Prerequisites
|
|
34
|
+
|
|
35
|
+
- Kubernetes cluster with GPU node pools (or cloud inference API access)
|
|
36
|
+
- Container registry (Harbor, ECR, GCR, or ACR)
|
|
37
|
+
- CI/CD system (GitHub Actions, GitLab CI, or Argo Workflows)
|
|
38
|
+
- Observability stack (Prometheus + Grafana + OpenTelemetry)
|
|
39
|
+
- Model registry (MLflow or custom metadata store)
|
|
40
|
+
|
|
41
|
+
## Outcomes
|
|
42
|
+
|
|
43
|
+
- Standardized path from experiment to production
|
|
44
|
+
- Safe model rollout with quality and safety gates
|
|
45
|
+
- Repeatable infra modules for inference, vector DB, and observability
|
|
46
|
+
- Clear ownership model across platform, app, and security teams
|
|
47
|
+
|
|
48
|
+
## Reference Architecture
|
|
49
|
+
|
|
50
|
+
1. **Control Plane**: model registry, prompt/version catalog, policy checks, eval pipeline.
|
|
51
|
+
2. **Data Plane**: inference gateway, vector database, cache, feature store.
|
|
52
|
+
3. **Ops Plane**: telemetry, alerting, SLO dashboards, cost analytics.
|
|
53
|
+
4. **Security Plane**: IAM boundaries, secret rotation, content filters, audit logs.
|
|
54
|
+
|
|
55
|
+
## Model Promotion Pipeline
|
|
56
|
+
|
|
57
|
+
```yaml
|
|
58
|
+
# .github/workflows/model-promotion.yaml
|
|
59
|
+
name: Model Promotion Pipeline
|
|
60
|
+
on:
|
|
61
|
+
workflow_dispatch:
|
|
62
|
+
inputs:
|
|
63
|
+
model_name:
|
|
64
|
+
description: "Model identifier"
|
|
65
|
+
required: true
|
|
66
|
+
model_version:
|
|
67
|
+
description: "Model version to promote"
|
|
68
|
+
required: true
|
|
69
|
+
target_env:
|
|
70
|
+
description: "Target environment"
|
|
71
|
+
required: true
|
|
72
|
+
type: choice
|
|
73
|
+
options: [staging, production]
|
|
74
|
+
|
|
75
|
+
jobs:
|
|
76
|
+
evaluate:
|
|
77
|
+
runs-on: ubuntu-latest
|
|
78
|
+
steps:
|
|
79
|
+
- uses: actions/checkout@v4
|
|
80
|
+
|
|
81
|
+
- name: Run quality evaluation suite
|
|
82
|
+
run: |
|
|
83
|
+
python -m evals.run \
|
|
84
|
+
--model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
|
|
85
|
+
--suite quality \
|
|
86
|
+
--output results/quality.json
|
|
87
|
+
|
|
88
|
+
- name: Run safety evaluation suite
|
|
89
|
+
run: |
|
|
90
|
+
python -m evals.run \
|
|
91
|
+
--model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
|
|
92
|
+
--suite safety \
|
|
93
|
+
--output results/safety.json
|
|
94
|
+
|
|
95
|
+
- name: Run latency benchmark
|
|
96
|
+
run: |
|
|
97
|
+
python -m evals.benchmark \
|
|
98
|
+
--model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
|
|
99
|
+
--concurrent-users 50 \
|
|
100
|
+
--duration 300 \
|
|
101
|
+
--output results/latency.json
|
|
102
|
+
|
|
103
|
+
- name: Gate check - quality
|
|
104
|
+
run: |
|
|
105
|
+
python -m evals.gate_check \
|
|
106
|
+
--results results/quality.json \
|
|
107
|
+
--threshold-file thresholds/quality.yaml
|
|
108
|
+
|
|
109
|
+
- name: Gate check - safety
|
|
110
|
+
run: |
|
|
111
|
+
python -m evals.gate_check \
|
|
112
|
+
--results results/safety.json \
|
|
113
|
+
--threshold-file thresholds/safety.yaml
|
|
114
|
+
|
|
115
|
+
- name: Gate check - latency
|
|
116
|
+
run: |
|
|
117
|
+
python -m evals.gate_check \
|
|
118
|
+
--results results/latency.json \
|
|
119
|
+
--threshold-file thresholds/latency.yaml
|
|
120
|
+
|
|
121
|
+
- name: Upload eval evidence
|
|
122
|
+
uses: actions/upload-artifact@v4
|
|
123
|
+
with:
|
|
124
|
+
name: eval-results-${{ inputs.model_version }}
|
|
125
|
+
path: results/
|
|
126
|
+
|
|
127
|
+
approve:
|
|
128
|
+
needs: evaluate
|
|
129
|
+
runs-on: ubuntu-latest
|
|
130
|
+
environment: ${{ inputs.target_env }}
|
|
131
|
+
steps:
|
|
132
|
+
- name: Record approval
|
|
133
|
+
run: |
|
|
134
|
+
echo "Approved by: ${{ github.actor }}"
|
|
135
|
+
echo "Model: ${{ inputs.model_name }}:${{ inputs.model_version }}"
|
|
136
|
+
echo "Target: ${{ inputs.target_env }}"
|
|
137
|
+
echo "Time: $(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
|
138
|
+
|
|
139
|
+
deploy:
|
|
140
|
+
needs: approve
|
|
141
|
+
runs-on: ubuntu-latest
|
|
142
|
+
steps:
|
|
143
|
+
- uses: actions/checkout@v4
|
|
144
|
+
|
|
145
|
+
- name: Deploy canary
|
|
146
|
+
run: |
|
|
147
|
+
kubectl set image deployment/${{ inputs.model_name }}-canary \
|
|
148
|
+
model=${{ inputs.model_name }}:${{ inputs.model_version }} \
|
|
149
|
+
-n ai-${{ inputs.target_env }}
|
|
150
|
+
|
|
151
|
+
- name: Wait for canary validation (15 min)
|
|
152
|
+
run: |
|
|
153
|
+
python -m canary.validate \
|
|
154
|
+
--deployment ${{ inputs.model_name }}-canary \
|
|
155
|
+
--namespace ai-${{ inputs.target_env }} \
|
|
156
|
+
--duration 900 \
|
|
157
|
+
--quality-threshold 0.85 \
|
|
158
|
+
--error-rate-threshold 0.02
|
|
159
|
+
|
|
160
|
+
- name: Promote to full rollout
|
|
161
|
+
run: |
|
|
162
|
+
kubectl set image deployment/${{ inputs.model_name }} \
|
|
163
|
+
model=${{ inputs.model_name }}:${{ inputs.model_version }} \
|
|
164
|
+
-n ai-${{ inputs.target_env }}
|
|
165
|
+
kubectl rollout status deployment/${{ inputs.model_name }} \
|
|
166
|
+
-n ai-${{ inputs.target_env }} --timeout=300s
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
## Evaluation Gate Thresholds
|
|
170
|
+
|
|
171
|
+
```yaml
|
|
172
|
+
# thresholds/quality.yaml
|
|
173
|
+
gates:
|
|
174
|
+
groundedness:
|
|
175
|
+
metric: groundedness_score
|
|
176
|
+
min: 0.85
|
|
177
|
+
comparison: gte
|
|
178
|
+
task_success:
|
|
179
|
+
metric: task_success_rate
|
|
180
|
+
min: 0.90
|
|
181
|
+
comparison: gte
|
|
182
|
+
hallucination:
|
|
183
|
+
metric: hallucination_rate
|
|
184
|
+
max: 0.08
|
|
185
|
+
comparison: lte
|
|
186
|
+
regression:
|
|
187
|
+
metric: quality_delta_vs_baseline
|
|
188
|
+
min: -0.02
|
|
189
|
+
comparison: gte
|
|
190
|
+
description: "Must not regress more than 2% vs current production"
|
|
191
|
+
|
|
192
|
+
# thresholds/latency.yaml
|
|
193
|
+
gates:
|
|
194
|
+
p50_latency:
|
|
195
|
+
metric: latency_p50_ms
|
|
196
|
+
max: 800
|
|
197
|
+
comparison: lte
|
|
198
|
+
p95_latency:
|
|
199
|
+
metric: latency_p95_ms
|
|
200
|
+
max: 2000
|
|
201
|
+
comparison: lte
|
|
202
|
+
p99_latency:
|
|
203
|
+
metric: latency_p99_ms
|
|
204
|
+
max: 5000
|
|
205
|
+
comparison: lte
|
|
206
|
+
throughput:
|
|
207
|
+
metric: requests_per_second
|
|
208
|
+
min: 50
|
|
209
|
+
comparison: gte
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
## A/B Testing Configuration
|
|
213
|
+
|
|
214
|
+
```yaml
|
|
215
|
+
# ab-test-config.yaml
|
|
216
|
+
apiVersion: gateway.ai/v1
|
|
217
|
+
kind: ABTest
|
|
218
|
+
metadata:
|
|
219
|
+
name: model-comparison-q1
|
|
220
|
+
namespace: ai-production
|
|
221
|
+
spec:
|
|
222
|
+
duration: 7d
|
|
223
|
+
traffic_split:
|
|
224
|
+
control:
|
|
225
|
+
model: gpt-4o-2024-08-06
|
|
226
|
+
weight: 70
|
|
227
|
+
treatment:
|
|
228
|
+
model: gpt-4o-2025-01-15
|
|
229
|
+
weight: 30
|
|
230
|
+
metrics:
|
|
231
|
+
primary:
|
|
232
|
+
- task_success_rate
|
|
233
|
+
- user_satisfaction_score
|
|
234
|
+
secondary:
|
|
235
|
+
- latency_p95
|
|
236
|
+
- cost_per_request
|
|
237
|
+
- hallucination_rate
|
|
238
|
+
guardrails:
|
|
239
|
+
auto_rollback_if:
|
|
240
|
+
- metric: task_success_rate
|
|
241
|
+
threshold: 0.80
|
|
242
|
+
window: 1h
|
|
243
|
+
- metric: hallucination_rate
|
|
244
|
+
threshold: 0.15
|
|
245
|
+
window: 30m
|
|
246
|
+
assignment:
|
|
247
|
+
strategy: sticky_user
|
|
248
|
+
hash_key: user_id
|
|
249
|
+
```
|
|
250
|
+
|
|
251
|
+
## Kubernetes Model Serving Deployment
|
|
252
|
+
|
|
253
|
+
```yaml
|
|
254
|
+
# model-serving-deployment.yaml
|
|
255
|
+
apiVersion: apps/v1
|
|
256
|
+
kind: Deployment
|
|
257
|
+
metadata:
|
|
258
|
+
name: llm-inference
|
|
259
|
+
namespace: ai-production
|
|
260
|
+
labels:
|
|
261
|
+
app: llm-inference
|
|
262
|
+
model: gpt-4o
|
|
263
|
+
version: "2025-01"
|
|
264
|
+
spec:
|
|
265
|
+
replicas: 3
|
|
266
|
+
strategy:
|
|
267
|
+
type: RollingUpdate
|
|
268
|
+
rollingUpdate:
|
|
269
|
+
maxSurge: 1
|
|
270
|
+
maxUnavailable: 0
|
|
271
|
+
selector:
|
|
272
|
+
matchLabels:
|
|
273
|
+
app: llm-inference
|
|
274
|
+
template:
|
|
275
|
+
metadata:
|
|
276
|
+
labels:
|
|
277
|
+
app: llm-inference
|
|
278
|
+
model: gpt-4o
|
|
279
|
+
annotations:
|
|
280
|
+
prometheus.io/scrape: "true"
|
|
281
|
+
prometheus.io/port: "8080"
|
|
282
|
+
prometheus.io/path: "/metrics"
|
|
283
|
+
spec:
|
|
284
|
+
topologySpreadConstraints:
|
|
285
|
+
- maxSkew: 1
|
|
286
|
+
topologyKey: topology.kubernetes.io/zone
|
|
287
|
+
whenUnsatisfiable: DoNotSchedule
|
|
288
|
+
labelSelector:
|
|
289
|
+
matchLabels:
|
|
290
|
+
app: llm-inference
|
|
291
|
+
containers:
|
|
292
|
+
- name: model
|
|
293
|
+
image: registry.internal/vllm-server:0.4.1
|
|
294
|
+
args:
|
|
295
|
+
- "--model=/models/current"
|
|
296
|
+
- "--tensor-parallel-size=1"
|
|
297
|
+
- "--max-model-len=8192"
|
|
298
|
+
- "--gpu-memory-utilization=0.90"
|
|
299
|
+
ports:
|
|
300
|
+
- containerPort: 8000
|
|
301
|
+
name: inference
|
|
302
|
+
- containerPort: 8080
|
|
303
|
+
name: metrics
|
|
304
|
+
resources:
|
|
305
|
+
requests:
|
|
306
|
+
cpu: "4"
|
|
307
|
+
memory: "16Gi"
|
|
308
|
+
nvidia.com/gpu: "1"
|
|
309
|
+
limits:
|
|
310
|
+
cpu: "8"
|
|
311
|
+
memory: "32Gi"
|
|
312
|
+
nvidia.com/gpu: "1"
|
|
313
|
+
readinessProbe:
|
|
314
|
+
httpGet:
|
|
315
|
+
path: /health
|
|
316
|
+
port: 8000
|
|
317
|
+
initialDelaySeconds: 60
|
|
318
|
+
periodSeconds: 10
|
|
319
|
+
livenessProbe:
|
|
320
|
+
httpGet:
|
|
321
|
+
path: /health
|
|
322
|
+
port: 8000
|
|
323
|
+
initialDelaySeconds: 120
|
|
324
|
+
periodSeconds: 30
|
|
325
|
+
volumeMounts:
|
|
326
|
+
- name: model-weights
|
|
327
|
+
mountPath: /models
|
|
328
|
+
readOnly: true
|
|
329
|
+
- name: config
|
|
330
|
+
mountPath: /etc/vllm
|
|
331
|
+
volumes:
|
|
332
|
+
- name: model-weights
|
|
333
|
+
persistentVolumeClaim:
|
|
334
|
+
claimName: model-weights-pvc
|
|
335
|
+
- name: config
|
|
336
|
+
configMap:
|
|
337
|
+
name: vllm-config
|
|
338
|
+
tolerations:
|
|
339
|
+
- key: nvidia.com/gpu
|
|
340
|
+
operator: Exists
|
|
341
|
+
effect: NoSchedule
|
|
342
|
+
nodeSelector:
|
|
343
|
+
gpu-type: a100
|
|
344
|
+
---
|
|
345
|
+
apiVersion: v1
|
|
346
|
+
kind: Service
|
|
347
|
+
metadata:
|
|
348
|
+
name: llm-inference
|
|
349
|
+
namespace: ai-production
|
|
350
|
+
spec:
|
|
351
|
+
selector:
|
|
352
|
+
app: llm-inference
|
|
353
|
+
ports:
|
|
354
|
+
- name: inference
|
|
355
|
+
port: 8000
|
|
356
|
+
targetPort: 8000
|
|
357
|
+
- name: metrics
|
|
358
|
+
port: 8080
|
|
359
|
+
targetPort: 8080
|
|
360
|
+
---
|
|
361
|
+
apiVersion: autoscaling/v2
|
|
362
|
+
kind: HorizontalPodAutoscaler
|
|
363
|
+
metadata:
|
|
364
|
+
name: llm-inference-hpa
|
|
365
|
+
namespace: ai-production
|
|
366
|
+
spec:
|
|
367
|
+
scaleTargetRef:
|
|
368
|
+
apiVersion: apps/v1
|
|
369
|
+
kind: Deployment
|
|
370
|
+
name: llm-inference
|
|
371
|
+
minReplicas: 2
|
|
372
|
+
maxReplicas: 10
|
|
373
|
+
metrics:
|
|
374
|
+
- type: Pods
|
|
375
|
+
pods:
|
|
376
|
+
metric:
|
|
377
|
+
name: llm_queue_depth
|
|
378
|
+
target:
|
|
379
|
+
type: AverageValue
|
|
380
|
+
averageValue: "5"
|
|
381
|
+
- type: Pods
|
|
382
|
+
pods:
|
|
383
|
+
metric:
|
|
384
|
+
name: gpu_utilization_percent
|
|
385
|
+
target:
|
|
386
|
+
type: AverageValue
|
|
387
|
+
averageValue: "75"
|
|
388
|
+
behavior:
|
|
389
|
+
scaleUp:
|
|
390
|
+
stabilizationWindowSeconds: 60
|
|
391
|
+
policies:
|
|
392
|
+
- type: Pods
|
|
393
|
+
value: 2
|
|
394
|
+
periodSeconds: 120
|
|
395
|
+
scaleDown:
|
|
396
|
+
stabilizationWindowSeconds: 300
|
|
397
|
+
policies:
|
|
398
|
+
- type: Pods
|
|
399
|
+
value: 1
|
|
400
|
+
periodSeconds: 300
|
|
401
|
+
```
|
|
402
|
+
|
|
403
|
+
## CI/CD Design for AI Services
|
|
404
|
+
|
|
405
|
+
- Build immutable containers with pinned dependencies and model hashes.
|
|
406
|
+
- Use environment promotion: `dev -> stage -> prod`.
|
|
407
|
+
- Fail deployment if:
|
|
408
|
+
- regression evals drop below baseline,
|
|
409
|
+
- safety tests exceed risk threshold,
|
|
410
|
+
- p95 latency exceeds SLO budget.
|
|
411
|
+
- Store deployment evidence for audits (commit SHA, eval report, approver).
|
|
412
|
+
|
|
413
|
+
## Operational SLOs
|
|
414
|
+
|
|
415
|
+
| Signal | Target | Measurement Window |
|
|
416
|
+
|--------|--------|--------------------|
|
|
417
|
+
| Availability | 99.9% | 30-day rolling |
|
|
418
|
+
| p95 Latency | < 1200ms | 5-min buckets |
|
|
419
|
+
| Cost per request | < $0.05 | 1-hour average |
|
|
420
|
+
| Task success rate | > 90% | 24-hour rolling |
|
|
421
|
+
| Groundedness | > 85% | 24-hour rolling |
|
|
422
|
+
|
|
423
|
+
## Platform Guardrails
|
|
424
|
+
|
|
425
|
+
- Enforce tenant quotas and model allow-lists.
|
|
426
|
+
- Require structured output contracts for automation paths.
|
|
427
|
+
- Default to low-risk model settings for critical workflows.
|
|
428
|
+
- Disable unconstrained tool execution in production.
|
|
429
|
+
|
|
430
|
+
## Tooling Stack (Example)
|
|
431
|
+
|
|
432
|
+
| Layer | Tools |
|
|
433
|
+
|-------|-------|
|
|
434
|
+
| Orchestration | Argo Workflows, GitHub Actions, Airflow |
|
|
435
|
+
| Model Registry | MLflow, custom metadata DB |
|
|
436
|
+
| Gateway | LiteLLM, Envoy-based API gateway |
|
|
437
|
+
| Observability | OpenTelemetry + Prometheus + Grafana + Langfuse |
|
|
438
|
+
| Policy | OPA/Rego for deployment and runtime checks |
|
|
439
|
+
| Evaluation | RAGAS, custom eval harness, Promptfoo |
|
|
440
|
+
| Serving | vLLM, TGI, Triton Inference Server |
|
|
441
|
+
|
|
442
|
+
## Troubleshooting
|
|
443
|
+
|
|
444
|
+
| Issue | Diagnosis | Resolution |
|
|
445
|
+
|-------|-----------|------------|
|
|
446
|
+
| Canary fails quality gate | Compare eval results with baseline | Adjust model config or revert version |
|
|
447
|
+
| Deployment stuck in rollout | Check pod events and resource quotas | Fix resource limits or node availability |
|
|
448
|
+
| A/B test shows no significant difference | Verify traffic split and sample size | Extend test duration or increase treatment weight |
|
|
449
|
+
| Model cold start too slow | Large model weight download | Use pre-cached PVCs or init containers |
|
|
450
|
+
| Eval pipeline flaky | Non-deterministic model outputs | Set temperature=0 for evals, increase sample size |
|
|
451
|
+
|
|
452
|
+
## Related Skills
|
|
453
|
+
|
|
454
|
+
- ai-pipeline-orchestration (`ai-pipeline-orchestration`) - Orchestrate ingestion and inference workflows
|
|
455
|
+
- agent-evals (`agent-evals`) - Build evaluation gates for releases
|
|
456
|
+
- llm-gateway (`llm-gateway`) - Route and control LLM traffic
|
|
457
|
+
- model-registry-governance (`model-registry-governance`) - Model lifecycle and approval workflows
|
|
458
|
+
- ai-sre-incident-response (`ai-sre-incident-response`) - AI-specific incident response
|
|
459
|
+
|
|
460
|
+
## Limitations
|
|
461
|
+
|
|
462
|
+
- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
|
|
463
|
+
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
|
|
464
|
+
|
|
465
|
+
### Example
|
|
466
|
+
|
|
467
|
+
```bash
|
|
468
|
+
git status && git diff --stat
|
|
469
|
+
kubectl diff -f manifest.yaml
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|