opencode-skills-collection 4.0.68 → 4.0.69
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +266 -1
- package/bundled-skills/access-review/SKILL.md +394 -0
- package/bundled-skills/access-review/references/details.md +121 -0
- package/bundled-skills/agent-evals/SKILL.md +420 -0
- package/bundled-skills/agent-observability/SKILL.md +346 -0
- package/bundled-skills/agent-observability/references/details.md +786 -0
- package/bundled-skills/ai-agent-security/SKILL.md +393 -0
- package/bundled-skills/ai-agent-security/references/details.md +912 -0
- package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
- package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
- package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
- package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
- package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
- package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
- package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
- package/bundled-skills/alerting-oncall/SKILL.md +458 -0
- package/bundled-skills/alerting-oncall/references/details.md +84 -0
- package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
- package/bundled-skills/argocd-gitops/SKILL.md +469 -0
- package/bundled-skills/arm-templates/SKILL.md +438 -0
- package/bundled-skills/arm-templates/references/details.md +64 -0
- package/bundled-skills/asset-inventory/SKILL.md +412 -0
- package/bundled-skills/asset-inventory/references/details.md +127 -0
- package/bundled-skills/audit-logging/SKILL.md +476 -0
- package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
- package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
- package/bundled-skills/aws-ec2/SKILL.md +426 -0
- package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
- package/bundled-skills/aws-iam/SKILL.md +463 -0
- package/bundled-skills/aws-lambda/SKILL.md +428 -0
- package/bundled-skills/aws-rds/SKILL.md +380 -0
- package/bundled-skills/aws-s3/SKILL.md +434 -0
- package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
- package/bundled-skills/aws-vpc/SKILL.md +436 -0
- package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
- package/bundled-skills/azure-aks/SKILL.md +423 -0
- package/bundled-skills/azure-devops/SKILL.md +457 -0
- package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
- package/bundled-skills/azure-keyvault/SKILL.md +455 -0
- package/bundled-skills/azure-keyvault/references/details.md +83 -0
- package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
- package/bundled-skills/azure-networking/SKILL.md +448 -0
- package/bundled-skills/azure-networking/references/details.md +135 -0
- package/bundled-skills/azure-sql/SKILL.md +413 -0
- package/bundled-skills/azure-sql/references/details.md +113 -0
- package/bundled-skills/azure-vms/SKILL.md +402 -0
- package/bundled-skills/azure-vms/references/details.md +134 -0
- package/bundled-skills/backup-recovery/SKILL.md +388 -0
- package/bundled-skills/bb-methodology/SKILL.md +451 -0
- package/bundled-skills/bb-methodology/references/details.md +120 -0
- package/bundled-skills/block-storage/SKILL.md +371 -0
- package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
- package/bundled-skills/blue-green-deploy/references/details.md +90 -0
- package/bundled-skills/bug-bounty/SKILL.md +447 -0
- package/bundled-skills/bug-bounty/references/details.md +1316 -0
- package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
- package/bundled-skills/business-continuity/SKILL.md +463 -0
- package/bundled-skills/career-ops/SKILL.md +186 -0
- package/bundled-skills/cdn-setup/SKILL.md +374 -0
- package/bundled-skills/change-management/SKILL.md +438 -0
- package/bundled-skills/change-management/references/details.md +105 -0
- package/bundled-skills/circleci/SKILL.md +475 -0
- package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
- package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
- package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
- package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
- package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
- package/bundled-skills/cloudformation/SKILL.md +461 -0
- package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
- package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
- package/bundled-skills/container-hardening/SKILL.md +126 -0
- package/bundled-skills/container-registries/SKILL.md +435 -0
- package/bundled-skills/container-scanning/SKILL.md +416 -0
- package/bundled-skills/convex-backend/SKILL.md +338 -0
- package/bundled-skills/dast-scanning/SKILL.md +437 -0
- package/bundled-skills/database-backups/SKILL.md +425 -0
- package/bundled-skills/datadog/SKILL.md +487 -0
- package/bundled-skills/dependency-scanning/SKILL.md +457 -0
- package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
- package/bundled-skills/disaster-recovery/SKILL.md +374 -0
- package/bundled-skills/disaster-recovery/references/details.md +219 -0
- package/bundled-skills/dns-management/SKILL.md +375 -0
- package/bundled-skills/docker-compose/SKILL.md +482 -0
- package/bundled-skills/docker-management/SKILL.md +426 -0
- package/bundled-skills/ebpf-observability/SKILL.md +436 -0
- package/bundled-skills/ebpf-observability/references/details.md +542 -0
- package/bundled-skills/elk-stack/SKILL.md +487 -0
- package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
- package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
- package/bundled-skills/feature-flags/SKILL.md +426 -0
- package/bundled-skills/feature-flags/references/details.md +86 -0
- package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
- package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
- package/bundled-skills/firewall-config/SKILL.md +479 -0
- package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
- package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
- package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
- package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
- package/bundled-skills/gcp-compute/SKILL.md +319 -0
- package/bundled-skills/gcp-gke/SKILL.md +307 -0
- package/bundled-skills/gcp-networking/SKILL.md +293 -0
- package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
- package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
- package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
- package/bundled-skills/gdpr-compliance/references/details.md +145 -0
- package/bundled-skills/geo-audit/SKILL.md +368 -0
- package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
- package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
- package/bundled-skills/geo-citability/SKILL.md +350 -0
- package/bundled-skills/geo-compare/SKILL.md +340 -0
- package/bundled-skills/geo-content/SKILL.md +383 -0
- package/bundled-skills/geo-crawlers/SKILL.md +408 -0
- package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
- package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
- package/bundled-skills/geo-proposal/SKILL.md +378 -0
- package/bundled-skills/geo-prospect/SKILL.md +225 -0
- package/bundled-skills/geo-report/SKILL.md +436 -0
- package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
- package/bundled-skills/geo-schema/SKILL.md +408 -0
- package/bundled-skills/geo-technical/SKILL.md +78 -0
- package/bundled-skills/geo-technical/references/details.md +543 -0
- package/bundled-skills/git-workflow/SKILL.md +460 -0
- package/bundled-skills/github-actions/SKILL.md +368 -0
- package/bundled-skills/gitlab-ci/SKILL.md +340 -0
- package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
- package/bundled-skills/gpu-server-management/SKILL.md +236 -0
- package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
- package/bundled-skills/helm-charts/SKILL.md +469 -0
- package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
- package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
- package/bundled-skills/hunt-ato/SKILL.md +184 -0
- package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
- package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
- package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
- package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
- package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
- package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
- package/bundled-skills/hunt-cicd/SKILL.md +311 -0
- package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
- package/bundled-skills/hunt-cors/SKILL.md +335 -0
- package/bundled-skills/hunt-dom/SKILL.md +323 -0
- package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
- package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
- package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
- package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
- package/bundled-skills/hunt-grpc/SKILL.md +317 -0
- package/bundled-skills/hunt-host-header/SKILL.md +309 -0
- package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
- package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
- package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
- package/bundled-skills/hunt-idor/SKILL.md +434 -0
- package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
- package/bundled-skills/hunt-k8s/SKILL.md +337 -0
- package/bundled-skills/hunt-laravel/SKILL.md +255 -0
- package/bundled-skills/hunt-ldap/SKILL.md +351 -0
- package/bundled-skills/hunt-lfi/SKILL.md +311 -0
- package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
- package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
- package/bundled-skills/hunt-misc/SKILL.md +378 -0
- package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
- package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
- package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
- package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
- package/bundled-skills/hunt-oauth/SKILL.md +459 -0
- package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
- package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
- package/bundled-skills/hunt-race-condition/references/details.md +159 -0
- package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
- package/bundled-skills/hunt-rce/SKILL.md +444 -0
- package/bundled-skills/hunt-rce/references/details.md +110 -0
- package/bundled-skills/hunt-saml/SKILL.md +156 -0
- package/bundled-skills/hunt-session/SKILL.md +342 -0
- package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
- package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
- package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
- package/bundled-skills/hunt-springboot/SKILL.md +285 -0
- package/bundled-skills/hunt-sqli/SKILL.md +466 -0
- package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
- package/bundled-skills/hunt-ssrf/references/details.md +179 -0
- package/bundled-skills/hunt-ssti/SKILL.md +163 -0
- package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
- package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
- package/bundled-skills/hunt-xxe/SKILL.md +466 -0
- package/bundled-skills/i-have-adhd/SKILL.md +170 -0
- package/bundled-skills/identity-access-management/SKILL.md +382 -0
- package/bundled-skills/identity-access-management/references/details.md +524 -0
- package/bundled-skills/incident-management/SKILL.md +484 -0
- package/bundled-skills/incident-response/SKILL.md +448 -0
- package/bundled-skills/incident-response/references/details.md +113 -0
- package/bundled-skills/interview-me/SKILL.md +248 -0
- package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
- package/bundled-skills/jenkins/SKILL.md +462 -0
- package/bundled-skills/jev-use/SKILL.md +158 -0
- package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
- package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
- package/bundled-skills/kubernetes-ops/references/details.md +108 -0
- package/bundled-skills/kustomize/SKILL.md +478 -0
- package/bundled-skills/linux-administration/SKILL.md +367 -0
- package/bundled-skills/linux-hardening/SKILL.md +154 -0
- package/bundled-skills/llm-app-security/SKILL.md +389 -0
- package/bundled-skills/llm-app-security/references/details.md +674 -0
- package/bundled-skills/llm-caching/SKILL.md +334 -0
- package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
- package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
- package/bundled-skills/llm-gateway/SKILL.md +282 -0
- package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
- package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
- package/bundled-skills/load-balancing/SKILL.md +403 -0
- package/bundled-skills/loki-logging/SKILL.md +479 -0
- package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
- package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
- package/bundled-skills/mcp-server-security/SKILL.md +356 -0
- package/bundled-skills/mcp-server-security/references/details.md +745 -0
- package/bundled-skills/mdm-device-management/SKILL.md +404 -0
- package/bundled-skills/mdm-device-management/references/details.md +410 -0
- package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
- package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
- package/bundled-skills/model-registry-governance/SKILL.md +452 -0
- package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
- package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
- package/bundled-skills/mongodb/SKILL.md +436 -0
- package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
- package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
- package/bundled-skills/mysql/SKILL.md +390 -0
- package/bundled-skills/new-relic/SKILL.md +472 -0
- package/bundled-skills/nfs-storage/SKILL.md +356 -0
- package/bundled-skills/object-storage/SKILL.md +378 -0
- package/bundled-skills/offensive-osint/SKILL.md +443 -0
- package/bundled-skills/okta-attack/SKILL.md +436 -0
- package/bundled-skills/ollama-stack/SKILL.md +379 -0
- package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
- package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
- package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
- package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
- package/bundled-skills/openshift/SKILL.md +485 -0
- package/bundled-skills/opentelemetry/SKILL.md +438 -0
- package/bundled-skills/opentelemetry/references/details.md +78 -0
- package/bundled-skills/opentofu-migration/SKILL.md +349 -0
- package/bundled-skills/osint-methodology/SKILL.md +460 -0
- package/bundled-skills/osint-methodology/references/details.md +1350 -0
- package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
- package/bundled-skills/penetration-testing/SKILL.md +152 -0
- package/bundled-skills/performance-tuning/SKILL.md +381 -0
- package/bundled-skills/planetscale/SKILL.md +297 -0
- package/bundled-skills/platform-engineering/SKILL.md +348 -0
- package/bundled-skills/platform-engineering/references/details.md +944 -0
- package/bundled-skills/podman/SKILL.md +405 -0
- package/bundled-skills/policy-as-code/SKILL.md +434 -0
- package/bundled-skills/policy-as-code/references/details.md +204 -0
- package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
- package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
- package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
- package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
- package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
- package/bundled-skills/rag-observability-evals/references/details.md +92 -0
- package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
- package/bundled-skills/redis/SKILL.md +421 -0
- package/bundled-skills/redteam-report-template/SKILL.md +370 -0
- package/bundled-skills/report-writing/SKILL.md +426 -0
- package/bundled-skills/report-writing/references/details.md +187 -0
- package/bundled-skills/reverse-proxy/SKILL.md +420 -0
- package/bundled-skills/runbook-creation/SKILL.md +438 -0
- package/bundled-skills/runbook-creation/references/details.md +71 -0
- package/bundled-skills/saas-security-posture/SKILL.md +415 -0
- package/bundled-skills/sast-scanning/SKILL.md +444 -0
- package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
- package/bundled-skills/security-arsenal/SKILL.md +446 -0
- package/bundled-skills/security-arsenal/references/details.md +540 -0
- package/bundled-skills/security-automation/SKILL.md +146 -0
- package/bundled-skills/semantic-versioning/SKILL.md +434 -0
- package/bundled-skills/semantic-versioning/references/details.md +83 -0
- package/bundled-skills/service-mesh/SKILL.md +422 -0
- package/bundled-skills/soc2-compliance/SKILL.md +409 -0
- package/bundled-skills/sops-encryption/SKILL.md +124 -0
- package/bundled-skills/sre-dashboards/SKILL.md +143 -0
- package/bundled-skills/ssh-configuration/SKILL.md +324 -0
- package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
- package/bundled-skills/ssl-tls-management/references/details.md +99 -0
- package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
- package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
- package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
- package/bundled-skills/systemd-services/SKILL.md +379 -0
- package/bundled-skills/terraform-aws/SKILL.md +125 -0
- package/bundled-skills/terraform-azure/SKILL.md +415 -0
- package/bundled-skills/terraform-azure/references/details.md +231 -0
- package/bundled-skills/terraform-gcp/SKILL.md +369 -0
- package/bundled-skills/threat-modeling/SKILL.md +487 -0
- package/bundled-skills/user-management/SKILL.md +383 -0
- package/bundled-skills/using-agent-skills/SKILL.md +220 -0
- package/bundled-skills/vector-database-ops/SKILL.md +300 -0
- package/bundled-skills/vendor-management/SKILL.md +439 -0
- package/bundled-skills/vendor-management/references/details.md +109 -0
- package/bundled-skills/vercel-deployments/SKILL.md +296 -0
- package/bundled-skills/vllm-server/SKILL.md +236 -0
- package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
- package/bundled-skills/vpn-setup/SKILL.md +452 -0
- package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
- package/bundled-skills/waf-setup/SKILL.md +354 -0
- package/bundled-skills/waf-setup/references/details.md +211 -0
- package/bundled-skills/web2-recon/SKILL.md +440 -0
- package/bundled-skills/web2-recon/references/details.md +319 -0
- package/bundled-skills/web3-audit/SKILL.md +445 -0
- package/bundled-skills/web3-audit/references/details.md +224 -0
- package/bundled-skills/windows-hardening/SKILL.md +454 -0
- package/bundled-skills/windows-hardening/references/details.md +204 -0
- package/bundled-skills/windows-server/SKILL.md +318 -0
- package/bundled-skills/zero-trust/SKILL.md +461 -0
- package/package.json +1 -1
- package/skills_index.json +6943 -323
|
@@ -0,0 +1,339 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: model-serving-kubernetes
|
|
3
|
+
description: Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA
|
|
4
|
+
Triton Inference Server.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant platform CLIs (kubectl, helm, terraform, git,
|
|
14
|
+
CI runners) and authorized access to the target environment. Docs-only; helper scripts
|
|
15
|
+
and templates not bundled.
|
|
16
|
+
metadata:
|
|
17
|
+
author: devops-skills
|
|
18
|
+
version: '1.0'
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
# Model Serving on Kubernetes
|
|
22
|
+
|
|
23
|
+
Production ML model serving with KServe and Triton — canary deployments, autoscaling, and GPU-aware scheduling.
|
|
24
|
+
|
|
25
|
+
## When to Use This Skill
|
|
26
|
+
|
|
27
|
+
Use this skill when:
|
|
28
|
+
- Serving scikit-learn, PyTorch, TensorFlow, or ONNX models at scale
|
|
29
|
+
- Implementing canary deployments and A/B testing for ML models
|
|
30
|
+
- Autoscaling inference pods based on request rate or GPU metrics
|
|
31
|
+
- Deploying LLMs with Triton or KServe on Kubernetes
|
|
32
|
+
- Managing multiple model versions with traffic splitting
|
|
33
|
+
|
|
34
|
+
## Prerequisites
|
|
35
|
+
|
|
36
|
+
- Kubernetes 1.28+ with GPU nodes
|
|
37
|
+
- KServe installed (or Triton standalone)
|
|
38
|
+
- `kubectl` and `helm` configured
|
|
39
|
+
- NVIDIA GPU Operator installed on cluster
|
|
40
|
+
|
|
41
|
+
## KServe Installation
|
|
42
|
+
|
|
43
|
+
```bash
|
|
44
|
+
# Install KServe with Helm
|
|
45
|
+
helm repo add kserve https://kserve.github.io/helm-charts
|
|
46
|
+
helm repo update
|
|
47
|
+
|
|
48
|
+
helm install kserve kserve/kserve \
|
|
49
|
+
--namespace kserve \
|
|
50
|
+
--create-namespace \
|
|
51
|
+
--set kserve.controller.gateway.ingressGateway.className=nginx
|
|
52
|
+
|
|
53
|
+
# Verify
|
|
54
|
+
kubectl get pods -n kserve
|
|
55
|
+
kubectl get crd | grep kserve
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
## Basic InferenceService (KServe)
|
|
59
|
+
|
|
60
|
+
```yaml
|
|
61
|
+
apiVersion: serving.kserve.io/v1beta1
|
|
62
|
+
kind: InferenceService
|
|
63
|
+
metadata:
|
|
64
|
+
name: sklearn-iris
|
|
65
|
+
namespace: models
|
|
66
|
+
spec:
|
|
67
|
+
predictor:
|
|
68
|
+
sklearn:
|
|
69
|
+
storageUri: gs://kfserving-examples/models/sklearn/1.0/model
|
|
70
|
+
resources:
|
|
71
|
+
requests:
|
|
72
|
+
cpu: "1"
|
|
73
|
+
memory: 2Gi
|
|
74
|
+
limits:
|
|
75
|
+
cpu: "2"
|
|
76
|
+
memory: 4Gi
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
kubectl apply -f inference-service.yaml
|
|
81
|
+
|
|
82
|
+
# Get inference service URL
|
|
83
|
+
kubectl get inferenceservice sklearn-iris -n models
|
|
84
|
+
# NAME URL READY ...
|
|
85
|
+
# sklearn-iris http://sklearn-iris.models.example.com True
|
|
86
|
+
|
|
87
|
+
# Test prediction
|
|
88
|
+
curl -X POST http://sklearn-iris.models.example.com/v1/models/sklearn-iris:predict \
|
|
89
|
+
-H "Content-Type: application/json" \
|
|
90
|
+
-d '{"instances": [[6.8, 2.8, 4.8, 1.4]]}'
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
## GPU-Enabled LLM InferenceService
|
|
94
|
+
|
|
95
|
+
```yaml
|
|
96
|
+
apiVersion: serving.kserve.io/v1beta1
|
|
97
|
+
kind: InferenceService
|
|
98
|
+
metadata:
|
|
99
|
+
name: llama-3-8b
|
|
100
|
+
namespace: models
|
|
101
|
+
annotations:
|
|
102
|
+
serving.kserve.io/enable-prometheus-scraping: "true"
|
|
103
|
+
spec:
|
|
104
|
+
predictor:
|
|
105
|
+
containers:
|
|
106
|
+
- name: vllm-container
|
|
107
|
+
image: vllm/vllm-openai:latest
|
|
108
|
+
args:
|
|
109
|
+
- "--model"
|
|
110
|
+
- "meta-llama/Llama-3.1-8B-Instruct"
|
|
111
|
+
- "--tensor-parallel-size"
|
|
112
|
+
- "1"
|
|
113
|
+
- "--gpu-memory-utilization"
|
|
114
|
+
- "0.90"
|
|
115
|
+
ports:
|
|
116
|
+
- containerPort: 8080
|
|
117
|
+
protocol: TCP
|
|
118
|
+
resources:
|
|
119
|
+
requests:
|
|
120
|
+
nvidia.com/gpu: "1"
|
|
121
|
+
memory: "20Gi"
|
|
122
|
+
cpu: "4"
|
|
123
|
+
limits:
|
|
124
|
+
nvidia.com/gpu: "1"
|
|
125
|
+
memory: "24Gi"
|
|
126
|
+
cpu: "8"
|
|
127
|
+
readinessProbe:
|
|
128
|
+
httpGet:
|
|
129
|
+
path: /health
|
|
130
|
+
port: 8080
|
|
131
|
+
initialDelaySeconds: 60
|
|
132
|
+
periodSeconds: 10
|
|
133
|
+
env:
|
|
134
|
+
- name: HUGGING_FACE_HUB_TOKEN
|
|
135
|
+
valueFrom:
|
|
136
|
+
secretKeyRef:
|
|
137
|
+
name: hf-token
|
|
138
|
+
key: token
|
|
139
|
+
nodeSelector:
|
|
140
|
+
nvidia.com/gpu.present: "true"
|
|
141
|
+
transformer:
|
|
142
|
+
containers:
|
|
143
|
+
- name: kserve-container
|
|
144
|
+
image: kserve/kserve-transformer:latest
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
## Canary Deployment (Traffic Splitting)
|
|
148
|
+
|
|
149
|
+
```yaml
|
|
150
|
+
apiVersion: serving.kserve.io/v1beta1
|
|
151
|
+
kind: InferenceService
|
|
152
|
+
metadata:
|
|
153
|
+
name: llama-3-8b
|
|
154
|
+
namespace: models
|
|
155
|
+
spec:
|
|
156
|
+
predictor:
|
|
157
|
+
canaryTrafficPercent: 20 # 20% to new version, 80% to stable
|
|
158
|
+
containers:
|
|
159
|
+
- name: vllm-container
|
|
160
|
+
image: vllm/vllm-openai:latest
|
|
161
|
+
args:
|
|
162
|
+
- "--model"
|
|
163
|
+
- "meta-llama/Llama-3.1-8B-Instruct-v2" # new model version
|
|
164
|
+
resources:
|
|
165
|
+
limits:
|
|
166
|
+
nvidia.com/gpu: "1"
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
```bash
|
|
170
|
+
# Gradually increase canary traffic
|
|
171
|
+
kubectl patch inferenceservice llama-3-8b -n models \
|
|
172
|
+
--type='json' \
|
|
173
|
+
-p='[{"op":"replace","path":"/spec/predictor/canaryTrafficPercent","value":50}]'
|
|
174
|
+
|
|
175
|
+
# Promote canary to stable
|
|
176
|
+
kubectl patch inferenceservice llama-3-8b -n models \
|
|
177
|
+
--type='json' \
|
|
178
|
+
-p='[{"op":"remove","path":"/spec/predictor/canaryTrafficPercent"}]'
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
## Autoscaling with KEDA
|
|
182
|
+
|
|
183
|
+
```yaml
|
|
184
|
+
apiVersion: keda.sh/v1alpha1
|
|
185
|
+
kind: ScaledObject
|
|
186
|
+
metadata:
|
|
187
|
+
name: llama-scaler
|
|
188
|
+
namespace: models
|
|
189
|
+
spec:
|
|
190
|
+
scaleTargetRef:
|
|
191
|
+
apiVersion: serving.kserve.io/v1beta1
|
|
192
|
+
kind: InferenceService
|
|
193
|
+
name: llama-3-8b
|
|
194
|
+
minReplicaCount: 1
|
|
195
|
+
maxReplicaCount: 5
|
|
196
|
+
triggers:
|
|
197
|
+
- type: prometheus
|
|
198
|
+
metadata:
|
|
199
|
+
serverAddress: http://prometheus-server.monitoring:9090
|
|
200
|
+
metricName: kserve_request_count
|
|
201
|
+
threshold: "10"
|
|
202
|
+
query: |
|
|
203
|
+
sum(rate(kserve_request_count_total{namespace="models",
|
|
204
|
+
service="llama-3-8b"}[1m]))
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
## NVIDIA Triton Inference Server
|
|
208
|
+
|
|
209
|
+
```yaml
|
|
210
|
+
apiVersion: apps/v1
|
|
211
|
+
kind: Deployment
|
|
212
|
+
metadata:
|
|
213
|
+
name: triton-server
|
|
214
|
+
namespace: models
|
|
215
|
+
spec:
|
|
216
|
+
replicas: 2
|
|
217
|
+
selector:
|
|
218
|
+
matchLabels:
|
|
219
|
+
app: triton
|
|
220
|
+
template:
|
|
221
|
+
metadata:
|
|
222
|
+
labels:
|
|
223
|
+
app: triton
|
|
224
|
+
spec:
|
|
225
|
+
containers:
|
|
226
|
+
- name: triton
|
|
227
|
+
image: nvcr.io/nvidia/tritonserver:24.05-py3
|
|
228
|
+
args:
|
|
229
|
+
- "tritonserver"
|
|
230
|
+
- "--model-store=s3://my-model-store/models"
|
|
231
|
+
- "--model-control-mode=poll" # auto-load new model versions
|
|
232
|
+
- "--repository-poll-secs=30"
|
|
233
|
+
- "--metrics-port=8002"
|
|
234
|
+
ports:
|
|
235
|
+
- containerPort: 8000 # HTTP
|
|
236
|
+
- containerPort: 8001 # gRPC
|
|
237
|
+
- containerPort: 8002 # Metrics
|
|
238
|
+
resources:
|
|
239
|
+
limits:
|
|
240
|
+
nvidia.com/gpu: "1"
|
|
241
|
+
readinessProbe:
|
|
242
|
+
httpGet:
|
|
243
|
+
path: /v2/health/ready
|
|
244
|
+
port: 8000
|
|
245
|
+
initialDelaySeconds: 30
|
|
246
|
+
```
|
|
247
|
+
|
|
248
|
+
## Triton Model Repository Structure
|
|
249
|
+
|
|
250
|
+
```
|
|
251
|
+
s3://my-model-store/models/
|
|
252
|
+
├── text-classifier/
|
|
253
|
+
│ ├── config.pbtxt
|
|
254
|
+
│ ├── 1/
|
|
255
|
+
│ │ └── model.onnx
|
|
256
|
+
│ └── 2/
|
|
257
|
+
│ └── model.onnx # new version; auto-loaded
|
|
258
|
+
├── embedding-model/
|
|
259
|
+
│ ├── config.pbtxt
|
|
260
|
+
│ └── 1/
|
|
261
|
+
│ └── model.onnx
|
|
262
|
+
```
|
|
263
|
+
|
|
264
|
+
```protobuf
|
|
265
|
+
# config.pbtxt for ONNX model
|
|
266
|
+
name: "text-classifier"
|
|
267
|
+
backend: "onnxruntime"
|
|
268
|
+
max_batch_size: 64
|
|
269
|
+
dynamic_batching {
|
|
270
|
+
preferred_batch_size: [16, 32]
|
|
271
|
+
max_queue_delay_microseconds: 1000
|
|
272
|
+
}
|
|
273
|
+
input [
|
|
274
|
+
{ name: "input_ids" data_type: TYPE_INT64 dims: [-1] }
|
|
275
|
+
{ name: "attention_mask" data_type: TYPE_INT64 dims: [-1] }
|
|
276
|
+
]
|
|
277
|
+
output [
|
|
278
|
+
{ name: "logits" data_type: TYPE_FP32 dims: [-1] }
|
|
279
|
+
]
|
|
280
|
+
instance_group [
|
|
281
|
+
{ kind: KIND_GPU count: 2 } # 2 model instances on GPU
|
|
282
|
+
]
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
## Model Management Commands
|
|
286
|
+
|
|
287
|
+
```bash
|
|
288
|
+
# List loaded models (Triton)
|
|
289
|
+
curl http://triton:8000/v2/models
|
|
290
|
+
|
|
291
|
+
# Load a new model version
|
|
292
|
+
curl -X POST http://triton:8000/v2/repository/models/text-classifier/load
|
|
293
|
+
|
|
294
|
+
# Unload a model
|
|
295
|
+
curl -X POST http://triton:8000/v2/repository/models/text-classifier/unload
|
|
296
|
+
|
|
297
|
+
# KServe — watch rollout status
|
|
298
|
+
kubectl rollout status deployment/llama-3-8b-predictor -n models
|
|
299
|
+
kubectl get inferenceservice llama-3-8b -n models -w
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
## Common Issues
|
|
303
|
+
|
|
304
|
+
| Issue | Cause | Fix |
|
|
305
|
+
|-------|-------|-----|
|
|
306
|
+
| `InferenceService not ready` | Model loading or OOM | Check predictor pod logs; increase memory limits |
|
|
307
|
+
| Canary stuck at 0% | KNative routing issue | Check `kubectl get ksvc -n models` |
|
|
308
|
+
| Triton missing model | S3 permissions or path | Verify IAM role; check `--model-store` path |
|
|
309
|
+
| Low GPU utilization | Dynamic batching off | Enable `dynamic_batching` in Triton config |
|
|
310
|
+
| Autoscaler not triggering | Prometheus query wrong | Test query in Prometheus UI |
|
|
311
|
+
|
|
312
|
+
## Best Practices
|
|
313
|
+
|
|
314
|
+
- Use canary deployments for all model updates — roll back in seconds if metrics degrade.
|
|
315
|
+
- Enable Triton dynamic batching — it can increase GPU throughput 5–10× for small models.
|
|
316
|
+
- Store models in S3/GCS with versioned paths (`s3://bucket/model/v1/`, `v2/`).
|
|
317
|
+
- Pin GPU node selectors to prevent model pods landing on CPU-only nodes.
|
|
318
|
+
- Monitor p99 latency and error rates per model version during canary rollouts.
|
|
319
|
+
|
|
320
|
+
## Related Skills
|
|
321
|
+
|
|
322
|
+
- vllm-server (`vllm-server`) - vLLM for LLM serving
|
|
323
|
+
- llm-inference-scaling (`llm-inference-scaling`) - KEDA autoscaling
|
|
324
|
+
- kubernetes-ops (`kubernetes-ops`) - Core Kubernetes operations
|
|
325
|
+
- gpu-server-management (`gpu-server-management`) - GPU nodes
|
|
326
|
+
|
|
327
|
+
## Limitations
|
|
328
|
+
|
|
329
|
+
- Guidance executes against real environments: confirm target, blast radius, and rollback plan before applying anything.
|
|
330
|
+
- Never deploy to production without explicit approval. Docs-only import: upstream scripts and templates not bundled.
|
|
331
|
+
|
|
332
|
+
### Example
|
|
333
|
+
|
|
334
|
+
```bash
|
|
335
|
+
git status && git diff --stat
|
|
336
|
+
kubectl diff -f manifest.yaml
|
|
337
|
+
```
|
|
338
|
+
|
|
339
|
+
> Adapted from [BagelHole/DevOps-Security-Agent-Skills](https://github.com/BagelHole/DevOps-Security-Agent-Skills) (MIT); frontmatter, When to Use/Limitations, and safety boundaries added for upstream compliance. Docs-only import: helper scripts and templates not bundled.
|