Dgx Spark Ops
Skill 62 of 183
Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark.
4 minutes · 958 words · 14 sections
Install
npx skills add wshobson/agents --skill spark-training-gotchasnpx skills add wshobson/agents/plugin marketplace add wshobson/agentsThe first command installs just this skill, by the name in its SKILL.md; the second installs the whole repository.
DGX Spark’s GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these checks. Read this before a long run, not after hour six.
nvidia-smi still shows headroom.| # | Symptom | Fix |
|---|---|---|
| G1 | undefined symbol / segfault | cu130 wheel or container |
| G2 | flash-attn wrong backend used | skip pip build; monkeypatch on NGC |
| G3 | OOM despite headroom | drop page cache |
| G4 | throughput drop / reboot | expect ~100W sustained cap |
| G5 | memory-bound step slow | budget 180–192 GB/s |
| G6 | cache evicted mid-run | one GPU server at a time |
| G7 | NVFP4 slower than FP8 | stay FP8 unless sm_121a |
| G8 | playbook fails outright | check upstream issues |
| G9 | env breaks after install | use a container |
| G10 | 2-Spark TP hangs | DDP/FSDP only, never TP |
ImportError: undefined symbol naming a CUDA
function, or a segfault on the first .cuda() call.libcudart.so.12; Spark
ships CUDA 13. pip never checks CUDA ABI, so it surfaces
only at import or first kernel launch.references/gotcha-checks.md G1 — the wheel’s
CUDA build tag.download.pytorch.org/whl/cu130 or
use a matched container.pip install flash-attn still fails/hangs.
Unsloth may also silently train flash-attn over an
explicitly requested SDPA.attn_implementation="sdpa".references/gotcha-checks.md G2 — is flash-attn
already present and working.references/gotcha-checks.md G2.nvidia-smi still reports free memory under the 128GB cap
— or, on some setups, [N/A] outright instead of a number.references/gotcha-checks.md G3 — read free -g
and /proc/meminfo, not nvidia-smi.sync; echo 3 > /proc/sys/vm/drop_caches — needs root, a
between-run reset, not a mid-training step.references/gotcha-checks.md G4 — sample
nvidia-smi --query-gpu=temperature.gpu,power.draw.references/gotcha-checks.md G5 — observed step
time vs. the measured range, not spec.gpu-memory-utilization<=0.5.references/gotcha-checks.md
G6 — other GPU-resident
processes and whether
capped.cvt.e2m1x2 unless kernels target
sm_121a; NVFP4 runs ~32% slower without it.references/gotcha-checks.md G7 — capability
reports (12, 1); does the build target sm_121a?sm_121a.references/gotcha-checks.md G8 — the playbook
repo’s recent issues.github.com/NVIDIA/dgx-spark-playbooks issues
before trusting a recipe for an expensive run.pip install, or two “identical”
environments behave differently.references/gotcha-checks.md
G9 — container or bare pip?spark-environment-setup
for tag guidance) or Unsloth’s container. If bare pip is
unavoidable, follow the NVIDIA install order, including
--no-deps on Unsloth.references/gotcha-checks.md G10 — the
configured parallelism strategy.The cheapest checks to run before anything else:
python3 -c "import torch; print(torch.version.cuda)" # expect 13.x (G1); NGC builds have no +cu130 tag — that's not a failureimport torch; print(torch.cuda.get_device_capability()) # expect (12, 1) (G7){ [ -f /.dockerenv -o -f /run/.containerenv ] || grep -qE 'docker|containerd' /proc/1/cgroup; } 2>/dev/null && echo container || echo unknown # G9assets/preflight.sh runs G1, G3, G4, G7, G9 and produces one
output line per gotcha in a fixed format: G-number first, then
PASS/FAIL/WARN where automatable, SKIP when unavailable, or
INFO: for a raw reading (G3, G4). Full commands:
references/gotcha-checks.md. See also
spark-environment-setup for the environment assumed working.
Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark. Use when a training run on DGX Spark fails to start, OOMs below the 128GB limit, slows down mid-run, or before any multi-hour training job on GB10.
The verbatim description from this skill’s front matter — the string an agent matches on to decide whether to load it.
main, last pushed 21 September 2026.SKILL.md, not by matching a directory convention. 51 distinct layouts observed: plugins/accessibility-compliance/skills/*/SKILL.md, plugins/agent-teams/skills/*/SKILL.md, plugins/api-scaffolding/skills/*/SKILL.md, plugins/avoid-ai-writing/skills/*/SKILL.md, plugins/backend-development/skills/*/SKILL.md, plugins/before-you-build/skills/*/SKILL.md, plugins/block-no-verify/skills/*/SKILL.md, plugins/blockchain-web3/skills/*/SKILL.md, plugins/brand-landingpage/skills/*/SKILL.md, plugins/business-analytics/skills/*/SKILL.md, plugins/cicd-automation/skills/*/SKILL.md, plugins/cloud-infrastructure/skills/*/SKILL.md, plugins/conductor/skills/*/SKILL.md, plugins/data-engineering/skills/*/SKILL.md, plugins/database-design/skills/*/SKILL.md, plugins/developer-essentials/skills/*/SKILL.md, plugins/dgx-spark-ops/skills/*/SKILL.md, plugins/documentation-generation/skills/*/SKILL.md, plugins/documentation-standards/skills/*/SKILL.md, plugins/dotnet-contribution/skills/*/SKILL.md, plugins/file-conversion/skills/*/SKILL.md, plugins/framework-migration/skills/*/SKILL.md, plugins/frontend-mobile-development/skills/*/SKILL.md, plugins/game-development/skills/*/SKILL.md, plugins/hermes-tweet/skills/*/SKILL.md, plugins/hr-legal-compliance/skills/*/SKILL.md, plugins/incident-response/skills/*/SKILL.md, plugins/javascript-typescript/skills/*/SKILL.md, plugins/kubernetes-operations/skills/*/SKILL.md, plugins/llm-application-dev/skills/*/SKILL.md, plugins/llm-finetuning/skills/*/SKILL.md, plugins/machine-learning-ops/skills/*/SKILL.md, plugins/observability-monitoring/skills/*/SKILL.md, plugins/payment-processing/skills/*/SKILL.md, plugins/plugin-eval/skills/*/SKILL.md, plugins/pptx-deck-creation/skills/*/SKILL.md, plugins/protect-mcp/skills/*/SKILL.md, plugins/python-development/skills/*/SKILL.md, plugins/quantitative-trading/skills/*/SKILL.md, plugins/reverse-engineering/skills/*/SKILL.md, plugins/review-agent-governance/skills/*/SKILL.md, plugins/security-scanning/skills/*/SKILL.md, plugins/shell-scripting/skills/*/SKILL.md, plugins/ship-mate/skills/*/SKILL.md, plugins/signed-audit-trails/skills/*/SKILL.md, plugins/skill-forge-essentials/skills/*/SKILL.md, plugins/social-publishing/skills/*/SKILL.md, plugins/startup-business-analyst/skills/*/SKILL.md, plugins/superself/skills/*/SKILL.md, plugins/systems-programming/skills/*/SKILL.md, plugins/ui-design/skills/*/SKILL.md.h1 and no skipped levels:.claude-plugin/marketplace.json by Seth Hobson, declaring 94 plugins. It is read for editorial metadata only — never as the skill index, which is always the repository tree./wshobson/agents.md, and each skill at its own .md URL.2 files · 12 KB
Everything this skill ships beside its prose. All of it is set here, as subchapters of skill 62.
Documentation the agent loads on demand, rather than up front.
Templates, schemas and fixtures the skill draws on.