opencode-skills-collection 4.0.68 → 4.0.70
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bundled-skills/.antigravity-install-manifest.json +304 -1
- package/bundled-skills/access-review/SKILL.md +394 -0
- package/bundled-skills/access-review/references/details.md +121 -0
- package/bundled-skills/agent-evals/SKILL.md +420 -0
- package/bundled-skills/agent-observability/SKILL.md +346 -0
- package/bundled-skills/agent-observability/references/details.md +786 -0
- package/bundled-skills/ai-agent-security/SKILL.md +393 -0
- package/bundled-skills/ai-agent-security/references/details.md +912 -0
- package/bundled-skills/ai-coding-agent-guardrails/SKILL.md +442 -0
- package/bundled-skills/ai-coding-agent-guardrails/references/details.md +753 -0
- package/bundled-skills/ai-inference-service-mesh/SKILL.md +449 -0
- package/bundled-skills/ai-pipeline-orchestration/SKILL.md +287 -0
- package/bundled-skills/ai-red-teaming/SKILL.md +409 -0
- package/bundled-skills/ai-security-hardening/SKILL.md +343 -0
- package/bundled-skills/ai-sre-incident-response/SKILL.md +336 -0
- package/bundled-skills/alerting-oncall/SKILL.md +458 -0
- package/bundled-skills/alerting-oncall/references/details.md +84 -0
- package/bundled-skills/api-integration-architect/SKILL.md +241 -0
- package/bundled-skills/apify-generate-output-schema/SKILL.md +438 -0
- package/bundled-skills/apify-integration-development/SKILL.md +168 -0
- package/bundled-skills/apify-integration-development/references/ai-framework-package.md +158 -0
- package/bundled-skills/apify-integration-development/references/ai-harness-plugin.md +192 -0
- package/bundled-skills/apify-integration-development/references/sdk-integration.md +236 -0
- package/bundled-skills/apify-integration-development/references/workflow-automation.md +163 -0
- package/bundled-skills/apk-redteam-pipeline/SKILL.md +446 -0
- package/bundled-skills/architecture-review/README.md +42 -0
- package/bundled-skills/architecture-review/SKILL.md +77 -0
- package/bundled-skills/architecture-review/examples.md +11 -0
- package/bundled-skills/architecture-review/reference/best-practices.md +7 -0
- package/bundled-skills/architecture-review/reference/capabilities.md +20 -0
- package/bundled-skills/architecture-review/reference/fallbacks.md +11 -0
- package/bundled-skills/architecture-review/reference/graph.md +15 -0
- package/bundled-skills/architecture-review/reference/mcp.md +14 -0
- package/bundled-skills/architecture-review/reference/workflow.md +15 -0
- package/bundled-skills/architecture-review/templates/architecture-review.md +21 -0
- package/bundled-skills/argocd-gitops/SKILL.md +469 -0
- package/bundled-skills/arm-templates/SKILL.md +438 -0
- package/bundled-skills/arm-templates/references/details.md +64 -0
- package/bundled-skills/asset-inventory/SKILL.md +412 -0
- package/bundled-skills/asset-inventory/references/details.md +127 -0
- package/bundled-skills/audit-logging/SKILL.md +476 -0
- package/bundled-skills/aws-cloudtrail/SKILL.md +486 -0
- package/bundled-skills/aws-cost-optimization/SKILL.md +331 -0
- package/bundled-skills/aws-ec2/SKILL.md +426 -0
- package/bundled-skills/aws-ecs-fargate/SKILL.md +388 -0
- package/bundled-skills/aws-iam/SKILL.md +463 -0
- package/bundled-skills/aws-lambda/SKILL.md +428 -0
- package/bundled-skills/aws-rds/SKILL.md +380 -0
- package/bundled-skills/aws-s3/SKILL.md +434 -0
- package/bundled-skills/aws-secrets-manager/SKILL.md +486 -0
- package/bundled-skills/aws-vpc/SKILL.md +436 -0
- package/bundled-skills/azure-ai-document-intelligence-ts/SKILL.md +1 -1
- package/bundled-skills/azure-aks/SKILL.md +423 -0
- package/bundled-skills/azure-devops/SKILL.md +457 -0
- package/bundled-skills/azure-functions-devsec/SKILL.md +436 -0
- package/bundled-skills/azure-keyvault/SKILL.md +455 -0
- package/bundled-skills/azure-keyvault/references/details.md +83 -0
- package/bundled-skills/azure-monitor-audit/SKILL.md +379 -0
- package/bundled-skills/azure-networking/SKILL.md +448 -0
- package/bundled-skills/azure-networking/references/details.md +135 -0
- package/bundled-skills/azure-sql/SKILL.md +413 -0
- package/bundled-skills/azure-sql/references/details.md +113 -0
- package/bundled-skills/azure-vms/SKILL.md +402 -0
- package/bundled-skills/azure-vms/references/details.md +134 -0
- package/bundled-skills/backup-recovery/SKILL.md +388 -0
- package/bundled-skills/bb-methodology/SKILL.md +451 -0
- package/bundled-skills/bb-methodology/references/details.md +120 -0
- package/bundled-skills/block-storage/SKILL.md +371 -0
- package/bundled-skills/blue-green-deploy/SKILL.md +453 -0
- package/bundled-skills/blue-green-deploy/references/details.md +90 -0
- package/bundled-skills/bug-bounty/SKILL.md +447 -0
- package/bundled-skills/bug-bounty/references/details.md +1316 -0
- package/bundled-skills/bugcrowd-reporting/SKILL.md +351 -0
- package/bundled-skills/business-continuity/SKILL.md +463 -0
- package/bundled-skills/career-ops/SKILL.md +186 -0
- package/bundled-skills/cdn-setup/SKILL.md +374 -0
- package/bundled-skills/change-management/SKILL.md +438 -0
- package/bundled-skills/change-management/references/details.md +105 -0
- package/bundled-skills/circleci/SKILL.md +475 -0
- package/bundled-skills/cis-benchmarks/SKILL.md +150 -0
- package/bundled-skills/cloudflare-pages/SKILL.md +318 -0
- package/bundled-skills/cloudflare-r2/SKILL.md +353 -0
- package/bundled-skills/cloudflare-workers/SKILL.md +415 -0
- package/bundled-skills/cloudflare-zero-trust/SKILL.md +361 -0
- package/bundled-skills/cloudformation/SKILL.md +461 -0
- package/bundled-skills/code-review-sensei/SKILL.md +177 -0
- package/bundled-skills/codebase-onboarding/README.md +42 -0
- package/bundled-skills/codebase-onboarding/SKILL.md +77 -0
- package/bundled-skills/codebase-onboarding/examples.md +11 -0
- package/bundled-skills/codebase-onboarding/reference/best-practices.md +7 -0
- package/bundled-skills/codebase-onboarding/reference/capabilities.md +20 -0
- package/bundled-skills/codebase-onboarding/reference/fallbacks.md +11 -0
- package/bundled-skills/codebase-onboarding/reference/graph.md +15 -0
- package/bundled-skills/codebase-onboarding/reference/mcp.md +14 -0
- package/bundled-skills/codebase-onboarding/reference/workflow.md +15 -0
- package/bundled-skills/codebase-onboarding/templates/repository-onboarding.md +21 -0
- package/bundled-skills/connection-auth-rules/SKILL.md +199 -0
- package/bundled-skills/connection-auth-rules/fetch_schema.py +320 -0
- package/bundled-skills/constraint-driven-development/SKILL.md +335 -0
- package/bundled-skills/constraint-driven-development/references/floor-guard.md +99 -0
- package/bundled-skills/container-hardening/SKILL.md +126 -0
- package/bundled-skills/container-registries/SKILL.md +435 -0
- package/bundled-skills/container-scanning/SKILL.md +416 -0
- package/bundled-skills/convex-backend/SKILL.md +338 -0
- package/bundled-skills/dast-scanning/SKILL.md +437 -0
- package/bundled-skills/database-backups/SKILL.md +425 -0
- package/bundled-skills/datadog/SKILL.md +487 -0
- package/bundled-skills/dependency-analysis/README.md +42 -0
- package/bundled-skills/dependency-analysis/SKILL.md +76 -0
- package/bundled-skills/dependency-analysis/examples.md +11 -0
- package/bundled-skills/dependency-analysis/reference/best-practices.md +7 -0
- package/bundled-skills/dependency-analysis/reference/capabilities.md +20 -0
- package/bundled-skills/dependency-analysis/reference/fallbacks.md +11 -0
- package/bundled-skills/dependency-analysis/reference/graph.md +15 -0
- package/bundled-skills/dependency-analysis/reference/mcp.md +14 -0
- package/bundled-skills/dependency-analysis/reference/workflow.md +15 -0
- package/bundled-skills/dependency-analysis/templates/dependency-review.md +21 -0
- package/bundled-skills/dependency-scanning/SKILL.md +457 -0
- package/bundled-skills/devcontainers-nix/SKILL.md +416 -0
- package/bundled-skills/devops-pipeline-builder/SKILL.md +200 -0
- package/bundled-skills/disaster-recovery/SKILL.md +374 -0
- package/bundled-skills/disaster-recovery/references/details.md +219 -0
- package/bundled-skills/dns-management/SKILL.md +375 -0
- package/bundled-skills/docker-compose/SKILL.md +482 -0
- package/bundled-skills/docker-management/SKILL.md +426 -0
- package/bundled-skills/eas-app-stores/SKILL.md +197 -0
- package/bundled-skills/eas-app-stores/agents/openai.yaml +4 -0
- package/bundled-skills/eas-app-stores/references/app-store-metadata.md +497 -0
- package/bundled-skills/eas-app-stores/references/ios-app-store.md +376 -0
- package/bundled-skills/eas-app-stores/references/native-ios.md +167 -0
- package/bundled-skills/eas-app-stores/references/play-store.md +244 -0
- package/bundled-skills/eas-app-stores/references/testflight.md +62 -0
- package/bundled-skills/eas-app-stores/references/workflows.md +120 -0
- package/bundled-skills/eas-hosting/SKILL.md +448 -0
- package/bundled-skills/eas-hosting/agents/openai.yaml +4 -0
- package/bundled-skills/eas-observe/SKILL.md +75 -0
- package/bundled-skills/eas-observe/agents/openai.yaml +4 -0
- package/bundled-skills/eas-observe/references/metrics.md +98 -0
- package/bundled-skills/eas-observe/references/queries.md +403 -0
- package/bundled-skills/eas-observe/references/setup.md +476 -0
- package/bundled-skills/eas-observe/references/third-party.md +136 -0
- package/bundled-skills/eas-simulator/SKILL.md +251 -0
- package/bundled-skills/eas-simulator/agents/openai.yaml +4 -0
- package/bundled-skills/eas-simulator/references/controllers.md +135 -0
- package/bundled-skills/eas-simulator/references/run-your-app.md +240 -0
- package/bundled-skills/eas-simulator/references/troubleshooting.md +47 -0
- package/bundled-skills/eas-workflows/SKILL.md +119 -0
- package/bundled-skills/eas-workflows/agents/openai.yaml +4 -0
- package/bundled-skills/eas-workflows/scripts/fetch.js +109 -0
- package/bundled-skills/ebpf-observability/SKILL.md +436 -0
- package/bundled-skills/ebpf-observability/references/details.md +542 -0
- package/bundled-skills/elk-stack/SKILL.md +487 -0
- package/bundled-skills/enterprise-vpn-attack/SKILL.md +395 -0
- package/bundled-skills/evidence-hygiene/SKILL.md +404 -0
- package/bundled-skills/expo-animation/LICENSE +21 -0
- package/bundled-skills/expo-animation/RECIPES.md +385 -0
- package/bundled-skills/expo-animation/SKILL.md +295 -0
- package/bundled-skills/expo-animation/agents/openai.yaml +4 -0
- package/bundled-skills/fact-check-x-unified/SKILL.md +178 -0
- package/bundled-skills/fact-check-x-unified/agents/openai.yaml +4 -0
- package/bundled-skills/fact-check-x-unified/references/acceptance-criteria.md +44 -0
- package/bundled-skills/fact-check-x-unified/references/contracts.md +39 -0
- package/bundled-skills/fact-check-x-unified/scripts/common.py +31 -0
- package/bundled-skills/fact-check-x-unified/scripts/fact_check_x.py +1832 -0
- package/bundled-skills/fact-check-x-unified/scripts/trusted_search_config.py +324 -0
- package/bundled-skills/fact-check-x-unified/tests/anchor_downgrade_test.py +90 -0
- package/bundled-skills/fact-check-x-unified/tests/multi_platform_test.py +369 -0
- package/bundled-skills/fact-check-x-unified/tests/smoke_test.py +740 -0
- package/bundled-skills/fact-check-x-unified/tests/stage_checkpoint_test.py +103 -0
- package/bundled-skills/fact-check-x-unified/tests/trusted_search_config_test.py +156 -0
- package/bundled-skills/feature-flags/SKILL.md +426 -0
- package/bundled-skills/feature-flags/references/details.md +86 -0
- package/bundled-skills/fedramp-compliance/SKILL.md +453 -0
- package/bundled-skills/firebase-app-platform/SKILL.md +381 -0
- package/bundled-skills/firewall-config/SKILL.md +479 -0
- package/bundled-skills/gcp-audit-logs/SKILL.md +452 -0
- package/bundled-skills/gcp-audit-logs/references/details.md +56 -0
- package/bundled-skills/gcp-cloud-functions/SKILL.md +284 -0
- package/bundled-skills/gcp-cloud-sql/SKILL.md +277 -0
- package/bundled-skills/gcp-compute/SKILL.md +319 -0
- package/bundled-skills/gcp-gke/SKILL.md +307 -0
- package/bundled-skills/gcp-networking/SKILL.md +293 -0
- package/bundled-skills/gcp-secret-manager/SKILL.md +421 -0
- package/bundled-skills/gcp-secret-manager/references/details.md +131 -0
- package/bundled-skills/gdpr-compliance/SKILL.md +451 -0
- package/bundled-skills/gdpr-compliance/references/details.md +145 -0
- package/bundled-skills/geo-audit/SKILL.md +368 -0
- package/bundled-skills/geo-brand-mentions/SKILL.md +68 -0
- package/bundled-skills/geo-brand-mentions/references/details.md +471 -0
- package/bundled-skills/geo-citability/SKILL.md +350 -0
- package/bundled-skills/geo-compare/SKILL.md +340 -0
- package/bundled-skills/geo-content/SKILL.md +383 -0
- package/bundled-skills/geo-crawlers/SKILL.md +408 -0
- package/bundled-skills/geo-llmstxt/SKILL.md +464 -0
- package/bundled-skills/geo-platform-optimizer/SKILL.md +314 -0
- package/bundled-skills/geo-proposal/SKILL.md +378 -0
- package/bundled-skills/geo-prospect/SKILL.md +225 -0
- package/bundled-skills/geo-report/SKILL.md +436 -0
- package/bundled-skills/geo-report-pdf/SKILL.md +157 -0
- package/bundled-skills/geo-schema/SKILL.md +408 -0
- package/bundled-skills/geo-technical/SKILL.md +78 -0
- package/bundled-skills/geo-technical/references/details.md +543 -0
- package/bundled-skills/git-workflow/SKILL.md +460 -0
- package/bundled-skills/github-actions/SKILL.md +368 -0
- package/bundled-skills/gitlab-ci/SKILL.md +340 -0
- package/bundled-skills/gpt-taste/SKILL.md +8 -1
- package/bundled-skills/gpu-kubernetes-operations/SKILL.md +468 -0
- package/bundled-skills/gpu-server-management/SKILL.md +236 -0
- package/bundled-skills/hashicorp-vault/SKILL.md +408 -0
- package/bundled-skills/helm-charts/SKILL.md +469 -0
- package/bundled-skills/hf-cli/SKILL.md +263 -0
- package/bundled-skills/hipaa-compliance/SKILL.md +451 -0
- package/bundled-skills/huggingface-community-evals/SKILL.md +228 -0
- package/bundled-skills/huggingface-community-evals/examples/.env.example +3 -0
- package/bundled-skills/huggingface-community-evals/examples/USAGE_EXAMPLES.md +101 -0
- package/bundled-skills/huggingface-community-evals/scripts/inspect_eval_uv.py +104 -0
- package/bundled-skills/huggingface-community-evals/scripts/inspect_vllm_uv.py +306 -0
- package/bundled-skills/huggingface-community-evals/scripts/lighteval_vllm_uv.py +297 -0
- package/bundled-skills/huggingface-datasets/SKILL.md +130 -0
- package/bundled-skills/hunt-aspnet/SKILL.md +321 -0
- package/bundled-skills/hunt-ato/SKILL.md +184 -0
- package/bundled-skills/hunt-auth-bypass/SKILL.md +426 -0
- package/bundled-skills/hunt-auth-bypass/references/details.md +80 -0
- package/bundled-skills/hunt-brute-force/SKILL.md +341 -0
- package/bundled-skills/hunt-business-logic/SKILL.md +281 -0
- package/bundled-skills/hunt-cache-poison/SKILL.md +382 -0
- package/bundled-skills/hunt-captcha-bypass/SKILL.md +136 -0
- package/bundled-skills/hunt-cicd/SKILL.md +311 -0
- package/bundled-skills/hunt-clickjacking/SKILL.md +110 -0
- package/bundled-skills/hunt-cors/SKILL.md +335 -0
- package/bundled-skills/hunt-dom/SKILL.md +323 -0
- package/bundled-skills/hunt-exceptional-conditions/SKILL.md +111 -0
- package/bundled-skills/hunt-file-upload/SKILL.md +202 -0
- package/bundled-skills/hunt-fintech-graphql/SKILL.md +289 -0
- package/bundled-skills/hunt-forgot-password/SKILL.md +114 -0
- package/bundled-skills/hunt-grpc/SKILL.md +317 -0
- package/bundled-skills/hunt-host-header/SKILL.md +309 -0
- package/bundled-skills/hunt-html-injection/SKILL.md +106 -0
- package/bundled-skills/hunt-http-smuggling/SKILL.md +129 -0
- package/bundled-skills/hunt-http-smuggling/references/phase2h-smuggling-cachepoison.md +177 -0
- package/bundled-skills/hunt-idor/SKILL.md +434 -0
- package/bundled-skills/hunt-jwt-crypto/SKILL.md +221 -0
- package/bundled-skills/hunt-k8s/SKILL.md +337 -0
- package/bundled-skills/hunt-laravel/SKILL.md +255 -0
- package/bundled-skills/hunt-ldap/SKILL.md +351 -0
- package/bundled-skills/hunt-lfi/SKILL.md +311 -0
- package/bundled-skills/hunt-llm-ai/SKILL.md +289 -0
- package/bundled-skills/hunt-mfa-bypass/SKILL.md +177 -0
- package/bundled-skills/hunt-misc/SKILL.md +378 -0
- package/bundled-skills/hunt-nextjs/SKILL.md +299 -0
- package/bundled-skills/hunt-nodejs/SKILL.md +263 -0
- package/bundled-skills/hunt-nosqli/SKILL.md +210 -0
- package/bundled-skills/hunt-ntlm-info/SKILL.md +314 -0
- package/bundled-skills/hunt-oauth/SKILL.md +459 -0
- package/bundled-skills/hunt-open-redirect/SKILL.md +223 -0
- package/bundled-skills/hunt-race-condition/SKILL.md +381 -0
- package/bundled-skills/hunt-race-condition/references/details.md +159 -0
- package/bundled-skills/hunt-rag-vector/SKILL.md +212 -0
- package/bundled-skills/hunt-rce/SKILL.md +444 -0
- package/bundled-skills/hunt-rce/references/details.md +110 -0
- package/bundled-skills/hunt-saml/SKILL.md +156 -0
- package/bundled-skills/hunt-session/SKILL.md +342 -0
- package/bundled-skills/hunt-shadow-api/SKILL.md +198 -0
- package/bundled-skills/hunt-source-leak/SKILL.md +345 -0
- package/bundled-skills/hunt-spa-api/SKILL.md +163 -0
- package/bundled-skills/hunt-springboot/SKILL.md +285 -0
- package/bundled-skills/hunt-sqli/SKILL.md +466 -0
- package/bundled-skills/hunt-ssrf/SKILL.md +396 -0
- package/bundled-skills/hunt-ssrf/references/details.md +179 -0
- package/bundled-skills/hunt-ssti/SKILL.md +163 -0
- package/bundled-skills/hunt-subdomain/SKILL.md +379 -0
- package/bundled-skills/hunt-tls-network/SKILL.md +399 -0
- package/bundled-skills/hunt-xxe/SKILL.md +466 -0
- package/bundled-skills/i-have-adhd/SKILL.md +170 -0
- package/bundled-skills/identity-access-management/SKILL.md +382 -0
- package/bundled-skills/identity-access-management/references/details.md +524 -0
- package/bundled-skills/incident-management/SKILL.md +484 -0
- package/bundled-skills/incident-response/SKILL.md +448 -0
- package/bundled-skills/incident-response/references/details.md +113 -0
- package/bundled-skills/interview-me/SKILL.md +248 -0
- package/bundled-skills/iso27001-compliance/SKILL.md +460 -0
- package/bundled-skills/jenkins/SKILL.md +462 -0
- package/bundled-skills/jev-social/SKILL.md +182 -0
- package/bundled-skills/jev-use/SKILL.md +158 -0
- package/bundled-skills/kubernetes-hardening/SKILL.md +154 -0
- package/bundled-skills/kubernetes-ops/SKILL.md +449 -0
- package/bundled-skills/kubernetes-ops/references/details.md +108 -0
- package/bundled-skills/kustomize/SKILL.md +478 -0
- package/bundled-skills/linux-administration/SKILL.md +367 -0
- package/bundled-skills/linux-hardening/SKILL.md +154 -0
- package/bundled-skills/llm-app-security/SKILL.md +389 -0
- package/bundled-skills/llm-app-security/references/details.md +674 -0
- package/bundled-skills/llm-caching/SKILL.md +334 -0
- package/bundled-skills/llm-cost-optimization/SKILL.md +311 -0
- package/bundled-skills/llm-fine-tuning/SKILL.md +329 -0
- package/bundled-skills/llm-gateway/SKILL.md +282 -0
- package/bundled-skills/llm-inference-scaling/SKILL.md +286 -0
- package/bundled-skills/llmops-platform-engineering/SKILL.md +472 -0
- package/bundled-skills/load-balancing/SKILL.md +403 -0
- package/bundled-skills/loki-logging/SKILL.md +479 -0
- package/bundled-skills/longbridge-derivatives/SKILL.md +117 -0
- package/bundled-skills/longbridge-derivatives/references/option.md +36 -0
- package/bundled-skills/longbridge-derivatives/references/options-advanced.md +101 -0
- package/bundled-skills/longbridge-derivatives/references/options-pnl.md +74 -0
- package/bundled-skills/longbridge-derivatives/references/options-strategy.md +82 -0
- package/bundled-skills/longbridge-derivatives/references/options-volatility.md +70 -0
- package/bundled-skills/longbridge-derivatives/references/warrant.md +12 -0
- package/bundled-skills/longbridge-quant/SKILL.md +151 -0
- package/bundled-skills/longbridge-quant/references/correlation.md +51 -0
- package/bundled-skills/longbridge-quant/references/execution-model.md +68 -0
- package/bundled-skills/longbridge-quant/references/factor-research.md +95 -0
- package/bundled-skills/longbridge-quant/references/factor-screen.md +101 -0
- package/bundled-skills/longbridge-quant/references/hedging.md +136 -0
- package/bundled-skills/longbridge-quant/references/ml-strategy.md +77 -0
- package/bundled-skills/longbridge-quant/references/multifactor.md +68 -0
- package/bundled-skills/longbridge-quant/references/pairs-trading.md +61 -0
- package/bundled-skills/longbridge-quant/references/quant-cli.md +133 -0
- package/bundled-skills/longbridge-quant/references/quant-stats.md +150 -0
- package/bundled-skills/longbridge-quant/references/seasonality.md +50 -0
- package/bundled-skills/longbridge-quant/references/strategy-optimizer.md +68 -0
- package/bundled-skills/longbridge-quant/references/volatility-strategy.md +52 -0
- package/bundled-skills/longbridge-research/SKILL.md +187 -0
- package/bundled-skills/longbridge-research/references/company-profile.md +96 -0
- package/bundled-skills/longbridge-research/references/company-tearsheet.md +82 -0
- package/bundled-skills/longbridge-research/references/competitive-analysis.md +81 -0
- package/bundled-skills/longbridge-research/references/consensus.md +92 -0
- package/bundled-skills/longbridge-research/references/coverage-initiation.md +76 -0
- package/bundled-skills/longbridge-research/references/defi-yield.md +60 -0
- package/bundled-skills/longbridge-research/references/finance-calendar.md +165 -0
- package/bundled-skills/longbridge-research/references/financial-planning.md +77 -0
- package/bundled-skills/longbridge-research/references/forecast-eps.md +39 -0
- package/bundled-skills/longbridge-research/references/fund-holder.md +44 -0
- package/bundled-skills/longbridge-research/references/hkipo-analysis.md +101 -0
- package/bundled-skills/longbridge-research/references/industry-peers.md +46 -0
- package/bundled-skills/longbridge-research/references/industry-rank.md +62 -0
- package/bundled-skills/longbridge-research/references/insider-trades.md +48 -0
- package/bundled-skills/longbridge-research/references/institution-rating.md +62 -0
- package/bundled-skills/longbridge-research/references/investment-ideas.md +69 -0
- package/bundled-skills/longbridge-research/references/investment-proposal.md +95 -0
- package/bundled-skills/longbridge-research/references/investors.md +87 -0
- package/bundled-skills/longbridge-research/references/onchain.md +70 -0
- package/bundled-skills/longbridge-research/references/post-investment.md +76 -0
- package/bundled-skills/longbridge-research/references/shareholder.md +72 -0
- package/bundled-skills/longbridge-research/references/short-positions.md +50 -0
- package/bundled-skills/longbridge-research/references/short-trades.md +50 -0
- package/bundled-skills/longbridge-research/references/stock-research.md +61 -0
- package/bundled-skills/longbridge-research/references/thesis-tracker.md +64 -0
- package/bundled-skills/m365-entra-attack/SKILL.md +423 -0
- package/bundled-skills/mac-mini-llm-lab/SKILL.md +350 -0
- package/bundled-skills/makepad-2-0-animation/SKILL.md +318 -0
- package/bundled-skills/makepad-2-0-animation/references/animator-reference.md +433 -0
- package/bundled-skills/makepad-2-0-dsl/SKILL.md +492 -0
- package/bundled-skills/makepad-2-0-dsl/references/dsl-syntax-reference.md +511 -0
- package/bundled-skills/makepad-2-0-dsl/references/extended-guide.md +56 -0
- package/bundled-skills/makepad-2-0-dsl/references/property-system.md +757 -0
- package/bundled-skills/makepad-2-0-events/SKILL.md +497 -0
- package/bundled-skills/makepad-2-0-events/references/event-patterns.md +802 -0
- package/bundled-skills/makepad-2-0-events/references/extended-guide.md +590 -0
- package/bundled-skills/makepad-2-0-layout/SKILL.md +499 -0
- package/bundled-skills/makepad-2-0-layout/references/extended-guide.md +243 -0
- package/bundled-skills/makepad-2-0-layout/references/layout-patterns.md +881 -0
- package/bundled-skills/makepad-2-0-widgets/SKILL.md +261 -0
- package/bundled-skills/makepad-2-0-widgets/references/widget-advanced.md +648 -0
- package/bundled-skills/makepad-2-0-widgets/references/widget-catalog.md +547 -0
- package/bundled-skills/mcp-server-security/SKILL.md +356 -0
- package/bundled-skills/mcp-server-security/references/details.md +745 -0
- package/bundled-skills/mdm-device-management/SKILL.md +404 -0
- package/bundled-skills/mdm-device-management/references/details.md +410 -0
- package/bundled-skills/meeting-distiller-pro/SKILL.md +120 -0
- package/bundled-skills/meme-coin-audit/SKILL.md +402 -0
- package/bundled-skills/mid-engagement-ir-detection/SKILL.md +377 -0
- package/bundled-skills/model-registry-governance/SKILL.md +452 -0
- package/bundled-skills/model-serving-kubernetes/SKILL.md +339 -0
- package/bundled-skills/model-supply-chain-security/SKILL.md +427 -0
- package/bundled-skills/mongodb/SKILL.md +436 -0
- package/bundled-skills/monte-carlo-analyze-root-cause/SKILL.md +12 -1
- package/bundled-skills/monte-carlo-asset-health/SKILL.md +12 -1
- package/bundled-skills/monte-carlo-context-detection/SKILL.md +170 -0
- package/bundled-skills/monte-carlo-context-detection/references/signal-definitions.md +46 -0
- package/bundled-skills/multi-tenant-llm-hosting/SKILL.md +435 -0
- package/bundled-skills/multi-tenant-llm-hosting/references/details.md +211 -0
- package/bundled-skills/mysql/SKILL.md +390 -0
- package/bundled-skills/new-relic/SKILL.md +472 -0
- package/bundled-skills/nfs-storage/SKILL.md +356 -0
- package/bundled-skills/object-storage/SKILL.md +378 -0
- package/bundled-skills/offensive-osint/SKILL.md +443 -0
- package/bundled-skills/okta-attack/SKILL.md +436 -0
- package/bundled-skills/ollama-stack/SKILL.md +379 -0
- package/bundled-skills/openclaw-deployment-hardening/SKILL.md +135 -0
- package/bundled-skills/openclaw-local-mac-mini/SKILL.md +426 -0
- package/bundled-skills/openclaw-local-mac-mini/references/details.md +221 -0
- package/bundled-skills/openclaw-security-hardening/SKILL.md +135 -0
- package/bundled-skills/openshift/SKILL.md +485 -0
- package/bundled-skills/opentelemetry/SKILL.md +438 -0
- package/bundled-skills/opentelemetry/references/details.md +78 -0
- package/bundled-skills/opentofu-migration/SKILL.md +349 -0
- package/bundled-skills/osint-methodology/SKILL.md +460 -0
- package/bundled-skills/osint-methodology/references/details.md +1350 -0
- package/bundled-skills/pci-dss-compliance/SKILL.md +446 -0
- package/bundled-skills/penetration-testing/SKILL.md +152 -0
- package/bundled-skills/performance-tuning/SKILL.md +381 -0
- package/bundled-skills/planetscale/SKILL.md +297 -0
- package/bundled-skills/platform-engineering/SKILL.md +348 -0
- package/bundled-skills/platform-engineering/references/details.md +944 -0
- package/bundled-skills/podman/SKILL.md +405 -0
- package/bundled-skills/policy-as-code/SKILL.md +434 -0
- package/bundled-skills/policy-as-code/references/details.md +204 -0
- package/bundled-skills/postgresql-devsec/SKILL.md +378 -0
- package/bundled-skills/prometheus-grafana/SKILL.md +469 -0
- package/bundled-skills/prompt-injection-defense/SKILL.md +483 -0
- package/bundled-skills/rag-infrastructure/SKILL.md +269 -0
- package/bundled-skills/rag-observability-evals/SKILL.md +444 -0
- package/bundled-skills/rag-observability-evals/references/details.md +92 -0
- package/bundled-skills/recon-scope-triage/SKILL.md +128 -0
- package/bundled-skills/redis/SKILL.md +421 -0
- package/bundled-skills/redteam-report-template/SKILL.md +370 -0
- package/bundled-skills/remotion-captions/SKILL.md +57 -0
- package/bundled-skills/remotion-captions/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-captions/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-captions/display-captions.md +190 -0
- package/bundled-skills/remotion-captions/import-srt-captions.md +73 -0
- package/bundled-skills/remotion-captions/transcribe-captions.md +70 -0
- package/bundled-skills/remotion-create/SKILL.md +106 -0
- package/bundled-skills/remotion-create/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-create/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-create/tailwind.md +11 -0
- package/bundled-skills/remotion-create/video-layout.md +9 -0
- package/bundled-skills/remotion-docs/SKILL.md +67 -0
- package/bundled-skills/remotion-docs/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-docs/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-interactivity/SKILL.md +270 -0
- package/bundled-skills/remotion-interactivity/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-interactivity/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-render/SKILL.md +48 -0
- package/bundled-skills/remotion-render/agents/openai.yaml +7 -0
- package/bundled-skills/remotion-render/assets/remotion-icon.svg +4 -0
- package/bundled-skills/remotion-render/transparent-videos.md +106 -0
- package/bundled-skills/report-writing/SKILL.md +426 -0
- package/bundled-skills/report-writing/references/details.md +187 -0
- package/bundled-skills/reverse-proxy/SKILL.md +420 -0
- package/bundled-skills/runbook-creation/SKILL.md +438 -0
- package/bundled-skills/runbook-creation/references/details.md +71 -0
- package/bundled-skills/saas-pricing-strategist/SKILL.md +169 -0
- package/bundled-skills/saas-security-posture/SKILL.md +415 -0
- package/bundled-skills/sast-scanning/SKILL.md +444 -0
- package/bundled-skills/sbom-supply-chain/SKILL.md +433 -0
- package/bundled-skills/score-eval/SKILL.md +35 -0
- package/bundled-skills/security-arsenal/SKILL.md +446 -0
- package/bundled-skills/security-arsenal/references/details.md +540 -0
- package/bundled-skills/security-automation/SKILL.md +146 -0
- package/bundled-skills/semantic-versioning/SKILL.md +434 -0
- package/bundled-skills/semantic-versioning/references/details.md +83 -0
- package/bundled-skills/service-mesh/SKILL.md +422 -0
- package/bundled-skills/soc2-compliance/SKILL.md +409 -0
- package/bundled-skills/sops-encryption/SKILL.md +124 -0
- package/bundled-skills/sre-dashboards/SKILL.md +143 -0
- package/bundled-skills/ssh-configuration/SKILL.md +324 -0
- package/bundled-skills/ssl-tls-management/SKILL.md +428 -0
- package/bundled-skills/ssl-tls-management/references/details.md +99 -0
- package/bundled-skills/startup-it-troubleshooting/SKILL.md +415 -0
- package/bundled-skills/supply-chain-attack-recon/SKILL.md +453 -0
- package/bundled-skills/supply-chain-attack-recon/references/details.md +258 -0
- package/bundled-skills/systemd-services/SKILL.md +379 -0
- package/bundled-skills/terraform-aws/SKILL.md +125 -0
- package/bundled-skills/terraform-azure/SKILL.md +415 -0
- package/bundled-skills/terraform-azure/references/details.md +231 -0
- package/bundled-skills/terraform-gcp/SKILL.md +369 -0
- package/bundled-skills/threat-modeling/SKILL.md +487 -0
- package/bundled-skills/user-management/SKILL.md +383 -0
- package/bundled-skills/using-agent-skills/SKILL.md +220 -0
- package/bundled-skills/vector-database-ops/SKILL.md +300 -0
- package/bundled-skills/vendor-management/SKILL.md +439 -0
- package/bundled-skills/vendor-management/references/details.md +109 -0
- package/bundled-skills/vercel-deployments/SKILL.md +296 -0
- package/bundled-skills/vllm-server/SKILL.md +236 -0
- package/bundled-skills/vmware-vcenter-attack/SKILL.md +412 -0
- package/bundled-skills/vpn-setup/SKILL.md +452 -0
- package/bundled-skills/vulnerability-scanning/SKILL.md +448 -0
- package/bundled-skills/waf-setup/SKILL.md +354 -0
- package/bundled-skills/waf-setup/references/details.md +211 -0
- package/bundled-skills/web2-recon/SKILL.md +440 -0
- package/bundled-skills/web2-recon/references/details.md +319 -0
- package/bundled-skills/web3-audit/SKILL.md +445 -0
- package/bundled-skills/web3-audit/references/details.md +224 -0
- package/bundled-skills/windows-hardening/SKILL.md +454 -0
- package/bundled-skills/windows-hardening/references/details.md +204 -0
- package/bundled-skills/windows-server/SKILL.md +318 -0
- package/bundled-skills/writing-guidelines/SKILL.md +60 -0
- package/bundled-skills/zero-trust/SKILL.md +461 -0
- package/package.json +1 -1
- package/skills_index.json +7874 -277
|
@@ -0,0 +1,468 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: gpu-kubernetes-operations
|
|
3
|
+
description: Operate GPU-backed Kubernetes clusters for AI inference and training
|
|
4
|
+
with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
|
|
5
|
+
category: devops
|
|
6
|
+
risk: critical
|
|
7
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
9
|
+
source_type: community
|
|
10
|
+
date_added: '2026-09-20'
|
|
11
|
+
license: MIT
|
|
12
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
13
|
+
compatibility: Requires the relevant OS/platform tooling and privileged access where
|
|
14
|
+
noted. Docs-only; helper scripts and templates not bundled.
|
|
15
|
+
metadata:
|
|
16
|
+
author: devops-skills
|
|
17
|
+
version: '1.0'
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
# GPU Kubernetes Operations
|
|
21
|
+
|
|
22
|
+
Run resilient and cost-efficient GPU clusters for production AI workloads.
|
|
23
|
+
|
|
24
|
+
## When to Use This Skill
|
|
25
|
+
|
|
26
|
+
- Setting up GPU node pools in Kubernetes for AI inference or training
|
|
27
|
+
- Configuring NVIDIA device plugin and GPU operator
|
|
28
|
+
- Implementing MIG partitioning to share GPUs across workloads
|
|
29
|
+
- Building GPU-aware autoscaling policies
|
|
30
|
+
- Monitoring GPU health with DCGM and Prometheus
|
|
31
|
+
- Troubleshooting GPU scheduling, driver, or OOM issues
|
|
32
|
+
|
|
33
|
+
## Prerequisites
|
|
34
|
+
|
|
35
|
+
- Kubernetes 1.28+ cluster with GPU-capable nodes
|
|
36
|
+
- NVIDIA GPUs (A10, L4, A100, H100, or similar)
|
|
37
|
+
- NVIDIA drivers installed on nodes (535+ recommended)
|
|
38
|
+
- Helm 3 for operator and plugin installation
|
|
39
|
+
- Prometheus stack for metrics collection
|
|
40
|
+
|
|
41
|
+
## NVIDIA GPU Operator Installation
|
|
42
|
+
|
|
43
|
+
The GPU Operator automates driver, toolkit, device plugin, and DCGM deployment.
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
# Add NVIDIA Helm repo
|
|
47
|
+
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
|
|
48
|
+
helm repo update
|
|
49
|
+
|
|
50
|
+
# Install GPU Operator
|
|
51
|
+
helm install gpu-operator nvidia/gpu-operator \
|
|
52
|
+
--namespace gpu-operator \
|
|
53
|
+
--create-namespace \
|
|
54
|
+
--set driver.enabled=true \
|
|
55
|
+
--set toolkit.enabled=true \
|
|
56
|
+
--set devicePlugin.enabled=true \
|
|
57
|
+
--set dcgmExporter.enabled=true \
|
|
58
|
+
--set migManager.enabled=true \
|
|
59
|
+
--set nodeStatusExporter.enabled=true \
|
|
60
|
+
--version v24.3.0
|
|
61
|
+
|
|
62
|
+
# Verify installation
|
|
63
|
+
kubectl get pods -n gpu-operator
|
|
64
|
+
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
## NVIDIA Device Plugin (Standalone)
|
|
68
|
+
|
|
69
|
+
If not using the GPU Operator, deploy the device plugin directly.
|
|
70
|
+
|
|
71
|
+
```yaml
|
|
72
|
+
# nvidia-device-plugin.yaml
|
|
73
|
+
apiVersion: apps/v1
|
|
74
|
+
kind: DaemonSet
|
|
75
|
+
metadata:
|
|
76
|
+
name: nvidia-device-plugin
|
|
77
|
+
namespace: kube-system
|
|
78
|
+
spec:
|
|
79
|
+
selector:
|
|
80
|
+
matchLabels:
|
|
81
|
+
name: nvidia-device-plugin
|
|
82
|
+
template:
|
|
83
|
+
metadata:
|
|
84
|
+
labels:
|
|
85
|
+
name: nvidia-device-plugin
|
|
86
|
+
spec:
|
|
87
|
+
tolerations:
|
|
88
|
+
- key: nvidia.com/gpu
|
|
89
|
+
operator: Exists
|
|
90
|
+
effect: NoSchedule
|
|
91
|
+
priorityClassName: system-node-critical
|
|
92
|
+
containers:
|
|
93
|
+
- name: nvidia-device-plugin
|
|
94
|
+
image: nvcr.io/nvidia/k8s-device-plugin:v0.15.0
|
|
95
|
+
securityContext:
|
|
96
|
+
privileged: true
|
|
97
|
+
env:
|
|
98
|
+
- name: FAIL_ON_INIT_ERROR
|
|
99
|
+
value: "false"
|
|
100
|
+
- name: DEVICE_SPLIT_COUNT
|
|
101
|
+
value: "1"
|
|
102
|
+
- name: DEVICE_LIST_STRATEGY
|
|
103
|
+
value: "envvar"
|
|
104
|
+
volumeMounts:
|
|
105
|
+
- name: device-plugin
|
|
106
|
+
mountPath: /var/lib/kubelet/device-plugins
|
|
107
|
+
volumes:
|
|
108
|
+
- name: device-plugin
|
|
109
|
+
hostPath:
|
|
110
|
+
path: /var/lib/kubelet/device-plugins
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
## MIG (Multi-Instance GPU) Partitioning
|
|
114
|
+
|
|
115
|
+
MIG allows a single A100 or H100 to be split into isolated GPU instances.
|
|
116
|
+
|
|
117
|
+
```yaml
|
|
118
|
+
# mig-config.yaml - ConfigMap for MIG Manager
|
|
119
|
+
apiVersion: v1
|
|
120
|
+
kind: ConfigMap
|
|
121
|
+
metadata:
|
|
122
|
+
name: mig-parted-config
|
|
123
|
+
namespace: gpu-operator
|
|
124
|
+
data:
|
|
125
|
+
config.yaml: |
|
|
126
|
+
version: v1
|
|
127
|
+
mig-configs:
|
|
128
|
+
# 7 small instances for inference microservices
|
|
129
|
+
all-1g.10gb:
|
|
130
|
+
- devices: all
|
|
131
|
+
mig-enabled: true
|
|
132
|
+
mig-devices:
|
|
133
|
+
"1g.10gb": 7
|
|
134
|
+
|
|
135
|
+
# 3 medium instances for mid-size models
|
|
136
|
+
all-2g.20gb:
|
|
137
|
+
- devices: all
|
|
138
|
+
mig-enabled: true
|
|
139
|
+
mig-devices:
|
|
140
|
+
"2g.20gb": 3
|
|
141
|
+
|
|
142
|
+
# Mixed: 1 large + 2 small
|
|
143
|
+
mixed-inference:
|
|
144
|
+
- devices: all
|
|
145
|
+
mig-enabled: true
|
|
146
|
+
mig-devices:
|
|
147
|
+
"3g.40gb": 1
|
|
148
|
+
"1g.10gb": 4
|
|
149
|
+
|
|
150
|
+
# Full GPU for training (no partitioning)
|
|
151
|
+
all-disabled:
|
|
152
|
+
- devices: all
|
|
153
|
+
mig-enabled: false
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
```bash
|
|
157
|
+
# Apply MIG profile to a node
|
|
158
|
+
kubectl label nodes gpu-node-01 nvidia.com/mig.config=all-1g.10gb --overwrite
|
|
159
|
+
|
|
160
|
+
# Verify MIG instances
|
|
161
|
+
kubectl exec -it nvidia-device-plugin-xxxxx -n kube-system -- nvidia-smi mig -lgi
|
|
162
|
+
|
|
163
|
+
# Check available MIG resources
|
|
164
|
+
kubectl get nodes gpu-node-01 -o json | jq '.status.allocatable | with_entries(select(.key | startswith("nvidia.com")))'
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
### Requesting MIG Slices in Pods
|
|
168
|
+
|
|
169
|
+
```yaml
|
|
170
|
+
# pod-with-mig.yaml
|
|
171
|
+
apiVersion: v1
|
|
172
|
+
kind: Pod
|
|
173
|
+
metadata:
|
|
174
|
+
name: inference-small
|
|
175
|
+
spec:
|
|
176
|
+
containers:
|
|
177
|
+
- name: model
|
|
178
|
+
image: registry.internal/vllm-server:latest
|
|
179
|
+
resources:
|
|
180
|
+
limits:
|
|
181
|
+
nvidia.com/mig-1g.10gb: 1
|
|
182
|
+
# For medium slice:
|
|
183
|
+
# nvidia.com/mig-2g.20gb: 1
|
|
184
|
+
# For large slice:
|
|
185
|
+
# nvidia.com/mig-3g.40gb: 1
|
|
186
|
+
```
|
|
187
|
+
|
|
188
|
+
## GPU Time-Slicing
|
|
189
|
+
|
|
190
|
+
For GPUs that do not support MIG (A10, L4), use time-slicing to share a GPU.
|
|
191
|
+
|
|
192
|
+
```yaml
|
|
193
|
+
# time-slicing-config.yaml
|
|
194
|
+
apiVersion: v1
|
|
195
|
+
kind: ConfigMap
|
|
196
|
+
metadata:
|
|
197
|
+
name: time-slicing-config
|
|
198
|
+
namespace: gpu-operator
|
|
199
|
+
data:
|
|
200
|
+
any: |-
|
|
201
|
+
version: v1
|
|
202
|
+
flags:
|
|
203
|
+
migStrategy: none
|
|
204
|
+
sharing:
|
|
205
|
+
timeSlicing:
|
|
206
|
+
renameByDefault: false
|
|
207
|
+
failRequestsGreaterThanOne: false
|
|
208
|
+
resources:
|
|
209
|
+
- name: nvidia.com/gpu
|
|
210
|
+
replicas: 4
|
|
211
|
+
```
|
|
212
|
+
|
|
213
|
+
```bash
|
|
214
|
+
# Apply time-slicing config
|
|
215
|
+
kubectl patch clusterpolicy/cluster-policy \
|
|
216
|
+
--type merge \
|
|
217
|
+
-p '{"spec":{"devicePlugin":{"config":{"name":"time-slicing-config","default":"any"}}}}'
|
|
218
|
+
|
|
219
|
+
# After applying, each physical GPU appears as 4 virtual GPUs
|
|
220
|
+
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'
|
|
221
|
+
# Output: "4" per physical GPU
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
## DCGM Monitoring
|
|
225
|
+
|
|
226
|
+
```yaml
|
|
227
|
+
# dcgm-servicemonitor.yaml
|
|
228
|
+
apiVersion: monitoring.coreos.com/v1
|
|
229
|
+
kind: ServiceMonitor
|
|
230
|
+
metadata:
|
|
231
|
+
name: dcgm-exporter
|
|
232
|
+
namespace: gpu-operator
|
|
233
|
+
labels:
|
|
234
|
+
release: prometheus
|
|
235
|
+
spec:
|
|
236
|
+
selector:
|
|
237
|
+
matchLabels:
|
|
238
|
+
app: nvidia-dcgm-exporter
|
|
239
|
+
endpoints:
|
|
240
|
+
- port: gpu-metrics
|
|
241
|
+
interval: 15s
|
|
242
|
+
path: /metrics
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
### Key DCGM Metrics and Alert Rules
|
|
246
|
+
|
|
247
|
+
```yaml
|
|
248
|
+
# gpu-alerts.yaml
|
|
249
|
+
groups:
|
|
250
|
+
- name: gpu-health
|
|
251
|
+
rules:
|
|
252
|
+
- alert: GPUHighTemperature
|
|
253
|
+
expr: DCGM_FI_DEV_GPU_TEMP > 85
|
|
254
|
+
for: 5m
|
|
255
|
+
labels:
|
|
256
|
+
severity: warning
|
|
257
|
+
annotations:
|
|
258
|
+
summary: "GPU {{ $labels.gpu }} temperature above 85C on {{ $labels.node }}"
|
|
259
|
+
|
|
260
|
+
- alert: GPUMemoryPressure
|
|
261
|
+
expr: (DCGM_FI_DEV_FB_USED / DCGM_FI_DEV_FB_FREE) > 0.90
|
|
262
|
+
for: 5m
|
|
263
|
+
labels:
|
|
264
|
+
severity: warning
|
|
265
|
+
annotations:
|
|
266
|
+
summary: "GPU memory above 90% on {{ $labels.node }} GPU {{ $labels.gpu }}"
|
|
267
|
+
|
|
268
|
+
- alert: GPUECCErrors
|
|
269
|
+
expr: increase(DCGM_FI_DEV_ECC_DBE_VOL_TOTAL[1h]) > 0
|
|
270
|
+
labels:
|
|
271
|
+
severity: critical
|
|
272
|
+
annotations:
|
|
273
|
+
summary: "Double-bit ECC errors detected on {{ $labels.node }} GPU {{ $labels.gpu }}"
|
|
274
|
+
|
|
275
|
+
- alert: GPUXidErrors
|
|
276
|
+
expr: increase(DCGM_FI_DEV_XID_ERRORS[5m]) > 0
|
|
277
|
+
labels:
|
|
278
|
+
severity: warning
|
|
279
|
+
annotations:
|
|
280
|
+
summary: "Xid error on {{ $labels.node }} GPU {{ $labels.gpu }}: {{ $labels.xid }}"
|
|
281
|
+
|
|
282
|
+
- alert: GPULowUtilization
|
|
283
|
+
expr: DCGM_FI_DEV_GPU_UTIL < 10 and on(pod) kube_pod_status_phase{phase="Running"} == 1
|
|
284
|
+
for: 30m
|
|
285
|
+
labels:
|
|
286
|
+
severity: info
|
|
287
|
+
annotations:
|
|
288
|
+
summary: "GPU underutilized on {{ $labels.node }} - consider rightsizing"
|
|
289
|
+
|
|
290
|
+
- alert: GPUDriverMismatch
|
|
291
|
+
expr: count(count by (driver_version)(DCGM_FI_DRIVER_VERSION)) > 1
|
|
292
|
+
labels:
|
|
293
|
+
severity: warning
|
|
294
|
+
annotations:
|
|
295
|
+
summary: "Multiple GPU driver versions detected across cluster"
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
## GPU Node Pool Configuration
|
|
299
|
+
|
|
300
|
+
```yaml
|
|
301
|
+
# gpu-nodepool.yaml
|
|
302
|
+
apiVersion: v1
|
|
303
|
+
kind: Node
|
|
304
|
+
metadata:
|
|
305
|
+
labels:
|
|
306
|
+
gpu-type: a100
|
|
307
|
+
gpu-memory: "80gb"
|
|
308
|
+
gpu-mig-capable: "true"
|
|
309
|
+
node-role: gpu-inference
|
|
310
|
+
spec:
|
|
311
|
+
taints:
|
|
312
|
+
- key: nvidia.com/gpu
|
|
313
|
+
value: "true"
|
|
314
|
+
effect: NoSchedule
|
|
315
|
+
---
|
|
316
|
+
# Inference deployment with GPU scheduling
|
|
317
|
+
apiVersion: apps/v1
|
|
318
|
+
kind: Deployment
|
|
319
|
+
metadata:
|
|
320
|
+
name: llm-inference
|
|
321
|
+
namespace: ai-serving
|
|
322
|
+
spec:
|
|
323
|
+
replicas: 3
|
|
324
|
+
selector:
|
|
325
|
+
matchLabels:
|
|
326
|
+
app: llm-inference
|
|
327
|
+
template:
|
|
328
|
+
metadata:
|
|
329
|
+
labels:
|
|
330
|
+
app: llm-inference
|
|
331
|
+
spec:
|
|
332
|
+
tolerations:
|
|
333
|
+
- key: nvidia.com/gpu
|
|
334
|
+
operator: Exists
|
|
335
|
+
effect: NoSchedule
|
|
336
|
+
nodeSelector:
|
|
337
|
+
gpu-type: a100
|
|
338
|
+
affinity:
|
|
339
|
+
podAntiAffinity:
|
|
340
|
+
preferredDuringSchedulingIgnoredDuringExecution:
|
|
341
|
+
- weight: 100
|
|
342
|
+
podAffinityTerm:
|
|
343
|
+
labelSelector:
|
|
344
|
+
matchLabels:
|
|
345
|
+
app: llm-inference
|
|
346
|
+
topologyKey: kubernetes.io/hostname
|
|
347
|
+
containers:
|
|
348
|
+
- name: vllm
|
|
349
|
+
image: registry.internal/vllm-server:0.4.1
|
|
350
|
+
resources:
|
|
351
|
+
requests:
|
|
352
|
+
nvidia.com/gpu: 1
|
|
353
|
+
cpu: "4"
|
|
354
|
+
memory: "32Gi"
|
|
355
|
+
limits:
|
|
356
|
+
nvidia.com/gpu: 1
|
|
357
|
+
cpu: "8"
|
|
358
|
+
memory: "64Gi"
|
|
359
|
+
env:
|
|
360
|
+
- name: CUDA_VISIBLE_DEVICES
|
|
361
|
+
value: "all"
|
|
362
|
+
```
|
|
363
|
+
|
|
364
|
+
## GPU Autoscaling
|
|
365
|
+
|
|
366
|
+
```yaml
|
|
367
|
+
# gpu-hpa.yaml
|
|
368
|
+
apiVersion: autoscaling/v2
|
|
369
|
+
kind: HorizontalPodAutoscaler
|
|
370
|
+
metadata:
|
|
371
|
+
name: llm-inference-hpa
|
|
372
|
+
namespace: ai-serving
|
|
373
|
+
spec:
|
|
374
|
+
scaleTargetRef:
|
|
375
|
+
apiVersion: apps/v1
|
|
376
|
+
kind: Deployment
|
|
377
|
+
name: llm-inference
|
|
378
|
+
minReplicas: 2
|
|
379
|
+
maxReplicas: 8
|
|
380
|
+
metrics:
|
|
381
|
+
- type: Pods
|
|
382
|
+
pods:
|
|
383
|
+
metric:
|
|
384
|
+
name: DCGM_FI_DEV_GPU_UTIL
|
|
385
|
+
target:
|
|
386
|
+
type: AverageValue
|
|
387
|
+
averageValue: "75"
|
|
388
|
+
- type: Pods
|
|
389
|
+
pods:
|
|
390
|
+
metric:
|
|
391
|
+
name: inference_queue_depth
|
|
392
|
+
target:
|
|
393
|
+
type: AverageValue
|
|
394
|
+
averageValue: "10"
|
|
395
|
+
behavior:
|
|
396
|
+
scaleUp:
|
|
397
|
+
stabilizationWindowSeconds: 60
|
|
398
|
+
policies:
|
|
399
|
+
- type: Pods
|
|
400
|
+
value: 2
|
|
401
|
+
periodSeconds: 120
|
|
402
|
+
scaleDown:
|
|
403
|
+
stabilizationWindowSeconds: 300
|
|
404
|
+
policies:
|
|
405
|
+
- type: Pods
|
|
406
|
+
value: 1
|
|
407
|
+
periodSeconds: 300
|
|
408
|
+
---
|
|
409
|
+
# Cluster Autoscaler config for GPU node pools
|
|
410
|
+
apiVersion: v1
|
|
411
|
+
kind: ConfigMap
|
|
412
|
+
metadata:
|
|
413
|
+
name: cluster-autoscaler-config
|
|
414
|
+
namespace: kube-system
|
|
415
|
+
data:
|
|
416
|
+
config: |
|
|
417
|
+
expander: priority
|
|
418
|
+
scale-down-delay-after-add: 10m
|
|
419
|
+
scale-down-unneeded-time: 10m
|
|
420
|
+
skip-nodes-with-local-storage: false
|
|
421
|
+
balance-similar-node-groups: true
|
|
422
|
+
expendable-pods-priority-cutoff: -10
|
|
423
|
+
gpu-total:
|
|
424
|
+
- min: 2
|
|
425
|
+
max: 16
|
|
426
|
+
gpu: nvidia.com/gpu
|
|
427
|
+
```
|
|
428
|
+
|
|
429
|
+
## Scheduling Patterns
|
|
430
|
+
|
|
431
|
+
- Use node affinity by GPU type (A10/L4/A100/H100).
|
|
432
|
+
- Separate latency-critical inference from batch training.
|
|
433
|
+
- Pin model replicas with anti-affinity for availability.
|
|
434
|
+
- Reserve headroom for failover and rolling updates.
|
|
435
|
+
|
|
436
|
+
## Cost Optimization
|
|
437
|
+
|
|
438
|
+
- Prefer MIG slices for smaller inference services.
|
|
439
|
+
- Schedule batch jobs in off-peak windows.
|
|
440
|
+
- Route low-priority traffic to cheaper model tiers.
|
|
441
|
+
- Use spot/preemptible instances for training workloads.
|
|
442
|
+
- Monitor GPU utilization and rightsize deployments.
|
|
443
|
+
|
|
444
|
+
## Troubleshooting
|
|
445
|
+
|
|
446
|
+
| Symptom | Check | Fix |
|
|
447
|
+
|---------|-------|-----|
|
|
448
|
+
| Pod stuck in Pending | `kubectl describe pod` for GPU resource events | Verify node has allocatable GPUs, check taints/tolerations |
|
|
449
|
+
| CUDA OOM during inference | Model too large for GPU memory | Reduce batch size, use quantization, or use MIG slice |
|
|
450
|
+
| DCGM metrics missing | ServiceMonitor labels matching | Verify DCGM exporter pod is running and scrape config |
|
|
451
|
+
| Driver mismatch after upgrade | `nvidia-smi` on each node | Cordon node, drain, upgrade driver, uncordon |
|
|
452
|
+
| GPU not detected | Device plugin pod logs | Restart device plugin, check NVIDIA container toolkit |
|
|
453
|
+
| Time-slicing not working | ConfigMap applied but no extra GPUs | Restart device plugin pods after config change |
|
|
454
|
+
| ECC errors increasing | `nvidia-smi -q -d ECC` | Schedule node drain and hardware replacement |
|
|
455
|
+
|
|
456
|
+
## Related Skills
|
|
457
|
+
|
|
458
|
+
- llm-inference-scaling (`llm-inference-scaling`) - Autoscale inference workloads
|
|
459
|
+
- model-serving-kubernetes (`model-serving-kubernetes`) - Production model serving patterns
|
|
460
|
+
- gpu-server-management (`gpu-server-management`) - Host-level GPU management fundamentals
|
|
461
|
+
- multi-tenant-llm-hosting (`multi-tenant-llm-hosting`) - Multi-tenant GPU sharing
|
|
462
|
+
- llm-cost-optimization (`llm-cost-optimization`) - Cost optimization strategies
|
|
463
|
+
|
|
464
|
+
## Limitations
|
|
465
|
+
|
|
466
|
+
- Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
|
|
467
|
+
- Docs-only import: upstream scripts and templates not bundled.
|
|
468
|
+
|
|
@@ -0,0 +1,236 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: gpu-server-management
|
|
3
|
+
description: Set up and manage NVIDIA GPU servers for AI workloads
|
|
4
|
+
category: devops
|
|
5
|
+
risk: critical
|
|
6
|
+
source: https://github.com/BagelHole/DevOps-Security-Agent-Skills
|
|
7
|
+
source_repo: BagelHole/DevOps-Security-Agent-Skills
|
|
8
|
+
source_type: community
|
|
9
|
+
date_added: '2026-09-20'
|
|
10
|
+
license: MIT
|
|
11
|
+
license_source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/blob/main/LICENSE
|
|
12
|
+
compatibility: Requires the relevant OS/platform tooling and privileged access where
|
|
13
|
+
noted. Docs-only; helper scripts and templates not bundled.
|
|
14
|
+
metadata:
|
|
15
|
+
author: devops-skills
|
|
16
|
+
version: '1.0'
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
# GPU Server Management
|
|
20
|
+
|
|
21
|
+
Provision, configure, and monitor NVIDIA GPU servers for AI inference and training workloads.
|
|
22
|
+
|
|
23
|
+
## When to Use This Skill
|
|
24
|
+
|
|
25
|
+
Use this skill when:
|
|
26
|
+
- Setting up a new GPU server for LLM inference or model training
|
|
27
|
+
- Installing or upgrading NVIDIA drivers and CUDA toolkit
|
|
28
|
+
- Configuring Docker with NVIDIA Container Toolkit for GPU workloads
|
|
29
|
+
- Partitioning A100/H100 GPUs with MIG for multi-tenant workloads
|
|
30
|
+
- Troubleshooting GPU errors, driver issues, or thermal throttling
|
|
31
|
+
|
|
32
|
+
## Prerequisites
|
|
33
|
+
|
|
34
|
+
- Ubuntu 22.04 LTS (recommended) or RHEL 8/9
|
|
35
|
+
- NVIDIA GPU (A10G, A100, H100, RTX 4090, or L40S recommended)
|
|
36
|
+
- Root or sudo access
|
|
37
|
+
- Internet access for package downloads
|
|
38
|
+
|
|
39
|
+
## Driver Installation (Ubuntu)
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
# Remove old drivers
|
|
43
|
+
sudo apt purge -y 'nvidia*' 'cuda*' 'libcuda*'
|
|
44
|
+
sudo apt autoremove -y
|
|
45
|
+
|
|
46
|
+
# Add NVIDIA package repository
|
|
47
|
+
distribution=$(. /etc/os-release; echo $ID$VERSION_ID)
|
|
48
|
+
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
|
|
49
|
+
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
|
|
50
|
+
|
|
51
|
+
curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
|
|
52
|
+
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
|
|
53
|
+
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
|
|
54
|
+
|
|
55
|
+
sudo apt update
|
|
56
|
+
|
|
57
|
+
# Install latest driver (560.x as of 2025)
|
|
58
|
+
sudo apt install -y nvidia-driver-560 cuda-toolkit-12-6
|
|
59
|
+
|
|
60
|
+
# Install NVIDIA Container Toolkit (Docker GPU support)
|
|
61
|
+
sudo apt install -y nvidia-container-toolkit
|
|
62
|
+
sudo nvidia-ctk runtime configure --runtime=docker
|
|
63
|
+
sudo systemctl restart docker
|
|
64
|
+
|
|
65
|
+
# Verify
|
|
66
|
+
nvidia-smi
|
|
67
|
+
nvcc --version
|
|
68
|
+
docker run --rm --gpus all nvidia/cuda:12.6.0-base-ubuntu22.04 nvidia-smi
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
## Post-Install Configuration
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
# Enable persistence mode (reduces driver initialization latency)
|
|
75
|
+
sudo nvidia-smi -pm 1
|
|
76
|
+
|
|
77
|
+
# Set power limit (reduce heat/noise on inference servers)
|
|
78
|
+
sudo nvidia-smi -pl 350 # watts; check TDP for your GPU model
|
|
79
|
+
|
|
80
|
+
# Disable ECC on inference servers (frees ~6% VRAM, less safe)
|
|
81
|
+
sudo nvidia-smi --ecc-config=0 # requires reboot
|
|
82
|
+
|
|
83
|
+
# Enable P2P for multi-GPU NVLink training
|
|
84
|
+
sudo nvidia-smi topo -m # check NVLink topology
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
## GPU Health Monitoring
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
# Real-time monitoring (like htop for GPUs)
|
|
91
|
+
watch -n 1 nvidia-smi
|
|
92
|
+
|
|
93
|
+
# Detailed stats
|
|
94
|
+
nvidia-smi --query-gpu=index,name,temperature.gpu,utilization.gpu,\
|
|
95
|
+
utilization.memory,memory.used,memory.free,power.draw,clocks.current.graphics \
|
|
96
|
+
--format=csv --loop=1
|
|
97
|
+
|
|
98
|
+
# DCGM — production monitoring daemon (for clusters)
|
|
99
|
+
sudo apt install -y datacenter-gpu-manager
|
|
100
|
+
sudo systemctl start dcgm
|
|
101
|
+
dcgmi discovery -l # list GPUs
|
|
102
|
+
dcgmi diag -r 1 # quick health check
|
|
103
|
+
dcgmi diag -r 3 # full diagnostic (takes ~20 min)
|
|
104
|
+
|
|
105
|
+
# Check GPU errors (XID errors — important for stability)
|
|
106
|
+
sudo dmesg | grep -i "NVRM\|nvidia\|XID"
|
|
107
|
+
nvidia-smi --query-gpu=ecc.errors.corrected.volatile.total \
|
|
108
|
+
--format=csv,noheader
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
## Prometheus GPU Metrics (DCGM Exporter)
|
|
112
|
+
|
|
113
|
+
```bash
|
|
114
|
+
# Deploy DCGM Exporter for Prometheus scraping
|
|
115
|
+
docker run -d \
|
|
116
|
+
--name dcgm-exporter \
|
|
117
|
+
--gpus all \
|
|
118
|
+
--cap-add SYS_ADMIN \
|
|
119
|
+
-p 9400:9400 \
|
|
120
|
+
--restart unless-stopped \
|
|
121
|
+
nvcr.io/nvidia/k8s/dcgm-exporter:latest
|
|
122
|
+
|
|
123
|
+
# Key metrics exposed:
|
|
124
|
+
# DCGM_FI_DEV_GPU_UTIL - GPU utilization %
|
|
125
|
+
# DCGM_FI_DEV_MEM_COPY_UTIL - Memory bandwidth utilization
|
|
126
|
+
# DCGM_FI_DEV_FB_USED - Framebuffer memory used (MB)
|
|
127
|
+
# DCGM_FI_DEV_SM_CLOCK - SM clock speed (MHz)
|
|
128
|
+
# DCGM_FI_DEV_GPU_TEMP - Temperature (°C)
|
|
129
|
+
# DCGM_FI_DEV_POWER_USAGE - Power draw (W)
|
|
130
|
+
# DCGM_FI_DEV_XID_ERRORS - XID error count (0 = healthy)
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
## MIG Partitioning (A100/H100)
|
|
134
|
+
|
|
135
|
+
MIG (Multi-Instance GPU) allows slicing one GPU into isolated smaller GPUs.
|
|
136
|
+
|
|
137
|
+
```bash
|
|
138
|
+
# Enable MIG mode (requires reboot or restart of all processes)
|
|
139
|
+
sudo nvidia-smi -mig 1
|
|
140
|
+
sudo systemctl restart nvidia-persistenced
|
|
141
|
+
|
|
142
|
+
# List available MIG profiles (A100 80GB example)
|
|
143
|
+
nvidia-smi mig -lgip
|
|
144
|
+
# 1g.10gb — 1 slice, 10GB (max 7 instances)
|
|
145
|
+
# 2g.20gb — 2 slices, 20GB (max 3 instances)
|
|
146
|
+
# 3g.40gb — 3 slices, 40GB (max 2 instances)
|
|
147
|
+
# 7g.80gb — full GPU, 80GB (max 1 instance)
|
|
148
|
+
|
|
149
|
+
# Create MIG instances (e.g., 3× 2g.20gb + 1× 2g.20gb = multi-tenant)
|
|
150
|
+
sudo nvidia-smi mig -cgi 2g.20gb,2g.20gb,2g.20gb,2g.20gb -C
|
|
151
|
+
|
|
152
|
+
# List created instances
|
|
153
|
+
nvidia-smi mig -lgi
|
|
154
|
+
nvidia-smi mig -lcgi
|
|
155
|
+
|
|
156
|
+
# Use in Docker
|
|
157
|
+
docker run --gpus '"device=MIG-GPU-xxx/0/0"' ...
|
|
158
|
+
|
|
159
|
+
# Disable MIG
|
|
160
|
+
sudo nvidia-smi mig -i 0 -dci
|
|
161
|
+
sudo nvidia-smi mig -i 0 -dgi
|
|
162
|
+
sudo nvidia-smi -mig 0
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
## Kernel & OS Tuning for GPU Servers
|
|
166
|
+
|
|
167
|
+
```bash
|
|
168
|
+
# Increase file descriptor limits
|
|
169
|
+
echo '* soft nofile 1048576' | sudo tee -a /etc/security/limits.conf
|
|
170
|
+
echo '* hard nofile 1048576' | sudo tee -a /etc/security/limits.conf
|
|
171
|
+
|
|
172
|
+
# Disable transparent huge pages (reduces latency jitter)
|
|
173
|
+
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
|
|
174
|
+
echo never | sudo tee /sys/kernel/mm/transparent_hugepage/defrag
|
|
175
|
+
|
|
176
|
+
# Persist via rc.local or systemd unit:
|
|
177
|
+
cat <<'EOF' | sudo tee /etc/rc.local
|
|
178
|
+
#!/bin/bash
|
|
179
|
+
echo never > /sys/kernel/mm/transparent_hugepage/enabled
|
|
180
|
+
echo never > /sys/kernel/mm/transparent_hugepage/defrag
|
|
181
|
+
nvidia-smi -pm 1
|
|
182
|
+
exit 0
|
|
183
|
+
EOF
|
|
184
|
+
sudo chmod +x /etc/rc.local
|
|
185
|
+
|
|
186
|
+
# PCIe performance mode
|
|
187
|
+
sudo nvidia-smi --auto-boost-default=0
|
|
188
|
+
sudo nvidia-smi --auto-boost-permission=0
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
## Multi-GPU Topology Check
|
|
192
|
+
|
|
193
|
+
```bash
|
|
194
|
+
# Check NVLink and PCIe topology
|
|
195
|
+
nvidia-smi topo -m
|
|
196
|
+
# Output shows interconnect type:
|
|
197
|
+
# NV4 = NVLink 4.0 (H100 SXM)
|
|
198
|
+
# NV2 = NVLink 2.0 (A100 SXM)
|
|
199
|
+
# PHB = PCIe bus (slower; avoid for tensor parallel training)
|
|
200
|
+
# PIX = same PCIe switch (fast)
|
|
201
|
+
|
|
202
|
+
# Bandwidth test between GPUs
|
|
203
|
+
/usr/local/cuda/samples/bin/x86_64/linux/release/p2pBandwidthLatencyTest
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
## Common Issues
|
|
207
|
+
|
|
208
|
+
| Issue | Cause | Fix |
|
|
209
|
+
|-------|-------|-----|
|
|
210
|
+
| `nvidia-smi: command not found` | Driver not installed | Follow driver installation steps above |
|
|
211
|
+
| Driver version mismatch | CUDA/driver incompatibility | Check compatibility matrix at developer.nvidia.com |
|
|
212
|
+
| GPU temperature >85°C | Poor airflow or fan failure | Check `nvidia-smi -q -d TEMPERATURE`; reseat cooler |
|
|
213
|
+
| XID 79 errors | GPU hardware error | Run `dcgmi diag -r 3`; may need GPU replacement |
|
|
214
|
+
| `failed to open device` in container | Container toolkit not configured | Run `nvidia-ctk runtime configure --runtime=docker` |
|
|
215
|
+
| Low PCIe bandwidth | Wrong slot or power limit | Check `nvidia-smi -q | grep PCIe`; use x16 slot |
|
|
216
|
+
|
|
217
|
+
## Best Practices
|
|
218
|
+
|
|
219
|
+
- Always enable persistence mode (`nvidia-smi -pm 1`) — reduces first-request latency.
|
|
220
|
+
- Monitor XID errors; persistent XID 79/94 indicates hardware failure.
|
|
221
|
+
- For training: use NVLink-connected GPUs; for inference: PCIe is usually fine.
|
|
222
|
+
- Set up DCGM alerts on temperature >80°C and power draw near TDP.
|
|
223
|
+
- Use MIG for multi-tenant inference to provide GPU isolation between models.
|
|
224
|
+
|
|
225
|
+
## Related Skills
|
|
226
|
+
|
|
227
|
+
- vllm-server (`vllm-server`) - LLM inference on GPUs
|
|
228
|
+
- llm-fine-tuning (`llm-fine-tuning`) - GPU training setup
|
|
229
|
+
- linux-hardening (`linux-hardening`) - Secure the host OS
|
|
230
|
+
- prometheus-grafana (`prometheus-grafana`) - Metrics dashboards
|
|
231
|
+
|
|
232
|
+
## Limitations
|
|
233
|
+
|
|
234
|
+
- Infrastructure commands can disrupt services: confirm target host/scope and have backups/snapshots before mutating state.
|
|
235
|
+
- Docs-only import: upstream scripts and templates not bundled.
|
|
236
|
+
|