art_pipe local dev: migrate to a CUDA 13 driver so default torch builds work #426

Closed
opened 2026-07-05 23:25:44 +00:00 by spikerj · 2 comments
Owner

Follow-up from #425 / spikersoft-backend PR #106 (Windows local ProArt worker).

Context

When bootstrapping the art_pipe model venvs, art_pipe's venv setup pulled a CUDA 13.0 torch build (torch 2.12.1+cu130), but the dev box's NVIDIA driver only advertises CUDA 12.9 (driver version 12090). Result: torch.cuda.is_available() was False and generation would fall back to CPU.

Current mitigation (shipped in PR #106)

scripts/windows-artpipe-bootstrap.ps1 pins torch to the cu128 index (CUDA 12.8 — the newest build compatible with a 12.9 driver, and the index the model manifests already declare). After each model bootstrap it checks CUDA availability in the venv and force-reinstalls torch/torchvision from cu128 if CUDA is unavailable. This keeps GPU generation working today without a host change.

The CUDA 13 path (this ticket)

We don't want to pin to cu128 forever. Decide + do the migration so the default (cu13x) torch builds work:

  • Upgrade the NVIDIA driver on the GPU dev/prod hosts to one that supports CUDA 13.x, then confirm a default pip install torch build reports cuda True.
  • Re-check the CUDA base image tag in SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile (currently nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04) and align it with the chosen toolkit.
  • Confirm the art_pipe model manifests' torch_index_url (currently cu128) — bump to cu13x once drivers are updated.
  • Once hosts are on a CUDA 13 driver, drop / relax the cu128 pin in windows-artpipe-bootstrap.ps1.
  • Verify the same holds for the production spikersoft-artpipe-* swarm stacks (SERVER GPU node driver version).

Acceptance

  • Default torch build reports cuda True on the target driver, no cu128 pin needed.
  • art_pipe stages (concept/modeling/texturing) run on GPU locally and in prod on the updated toolchain.
Follow-up from #425 / spikersoft-backend PR #106 (Windows local ProArt worker). ## Context When bootstrapping the art_pipe model venvs, art_pipe's venv setup pulled a **CUDA 13.0** torch build (`torch 2.12.1+cu130`), but the dev box's NVIDIA driver only advertises **CUDA 12.9** (driver version 12090). Result: `torch.cuda.is_available()` was `False` and generation would fall back to CPU. ## Current mitigation (shipped in PR #106) `scripts/windows-artpipe-bootstrap.ps1` pins torch to the **cu128** index (CUDA 12.8 — the newest build compatible with a 12.9 driver, and the index the model manifests already declare). After each model bootstrap it checks CUDA availability in the venv and force-reinstalls torch/torchvision from cu128 if CUDA is unavailable. This keeps GPU generation working today without a host change. ## The CUDA 13 path (this ticket) We don't want to pin to cu128 forever. Decide + do the migration so the default (cu13x) torch builds work: - Upgrade the NVIDIA driver on the GPU dev/prod hosts to one that supports CUDA 13.x, then confirm a default `pip install torch` build reports `cuda True`. - Re-check the CUDA base image tag in `SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile` (currently `nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04`) and align it with the chosen toolkit. - Confirm the art_pipe model manifests' `torch_index_url` (currently cu128) — bump to cu13x once drivers are updated. - Once hosts are on a CUDA 13 driver, drop / relax the cu128 pin in `windows-artpipe-bootstrap.ps1`. - Verify the same holds for the production spikersoft-artpipe-* swarm stacks (SERVER GPU node driver version). ## Acceptance - Default torch build reports `cuda True` on the target driver, no cu128 pin needed. - art_pipe stages (concept/modeling/texturing) run on GPU locally and in prod on the updated toolchain.
Author
Owner

Board-sweep status (2026-07-22): mitigation shipped (cu128 torch pin in the bootstrap script, backend #106). REMAINING: the actual dev-box CUDA-13 driver migration (ops).

Board-sweep status (2026-07-22): mitigation shipped (cu128 torch pin in the bootstrap script, backend #106). REMAINING: the actual dev-box CUDA-13 driver migration (ops).
Author
Owner

Migrated as part of the umbrella-tracker breakup.

This ticket needed changes in more than one repo, so it became one issue per repo:

The shape of this ticket has changed and both children record it: the driver half is done.
nvidia-smi on the GPU host reports NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0 on the RTX 4090. So the ops step this was blocked on has happened, and what remains is purely the code still pinned to the old toolchain.

Verified 2026-08-07: on spikersoft-artpipe@7082d26, grep -rho "cu1[0-9]*" models/*/artpipe.json returns cu128 x 20 — every manifest is still pinned. On spikersoft-backend@98102023, SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile:6 is still nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04 and the ps1 still carries the cu128 pin.

Recorded on both children as a precondition: confirm the SERVER lane also reports CUDA 13.x before bumping — the coordinator now runs two enabled capacity nodes, and #822 left an unverified NVML driver/library mismatch on SERVER. Bumping while SERVER is on 12.x would move the breakage rather than fix it.

Status: partially done — mitigation shipped (backend #106) and the driver upgraded on the 4090; the manifest bump, the base-image bump and the pin removal remain.

Closing here. Work now lives in the repos that hold the fixes, so fixes #N in a PR will auto-close them on merge. The umbrella tracker keeps cross-repo epics only.

— Opus 5 Agent

Migrated as part of the umbrella-tracker breakup. This ticket needed changes in more than one repo, so it became one issue per repo: - ArtPipe: spikerj/spikersoft-artpipe#59 — bump the model manifests' `env.torch_index_url` off cu128 and re-verify the venvs. - Backend: spikerj/spikersoft-backend#608 — align the ArtPipeProcessor CUDA base image and drop the cu128 pin from `windows-artpipe-bootstrap.ps1`. **The shape of this ticket has changed and both children record it: the driver half is done.** `nvidia-smi` on the GPU host reports `NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0` on the RTX 4090. So the ops step this was blocked on has happened, and what remains is purely the code still pinned to the old toolchain. Verified 2026-08-07: on `spikersoft-artpipe@7082d26`, `grep -rho "cu1[0-9]*" models/*/artpipe.json` returns cu128 x 20 — every manifest is still pinned. On `spikersoft-backend@98102023`, `SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile:6` is still `nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04` and the ps1 still carries the cu128 pin. Recorded on both children as a precondition: confirm the SERVER lane also reports CUDA 13.x before bumping — the coordinator now runs two enabled capacity nodes, and #822 left an unverified NVML driver/library mismatch on SERVER. Bumping while SERVER is on 12.x would move the breakage rather than fix it. Status: partially done — mitigation shipped (backend #106) and the driver upgraded on the 4090; the manifest bump, the base-image bump and the pin removal remain. Closing here. Work now lives in the repos that hold the fixes, so `fixes #N` in a PR will auto-close them on merge. The umbrella tracker keeps cross-repo epics only. — Opus 5 Agent
Sign in to join this conversation.