Follow-up from #425 / spikersoft-backend PR #106 (Windows local ProArt worker).
Context
When bootstrapping the art_pipe model venvs, art_pipe's venv setup pulled a CUDA 13.0 torch build (torch 2.12.1+cu130), but the dev box's NVIDIA driver only advertises CUDA 12.9 (driver version 12090). Result: torch.cuda.is_available() was False and generation would fall back to CPU.
scripts/windows-artpipe-bootstrap.ps1 pins torch to the cu128 index (CUDA 12.8 — the newest build compatible with a 12.9 driver, and the index the model manifests already declare). After each model bootstrap it checks CUDA availability in the venv and force-reinstalls torch/torchvision from cu128 if CUDA is unavailable. This keeps GPU generation working today without a host change.
The CUDA 13 path (this ticket)
We don't want to pin to cu128 forever. Decide + do the migration so the default (cu13x) torch builds work:
Upgrade the NVIDIA driver on the GPU dev/prod hosts to one that supports CUDA 13.x, then confirm a default pip install torch build reports cuda True.
Re-check the CUDA base image tag in SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile (currently nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04) and align it with the chosen toolkit.
Confirm the art_pipe model manifests' torch_index_url (currently cu128) — bump to cu13x once drivers are updated.
Once hosts are on a CUDA 13 driver, drop / relax the cu128 pin in windows-artpipe-bootstrap.ps1.
Verify the same holds for the production spikersoft-artpipe-* swarm stacks (SERVER GPU node driver version).
Acceptance
Default torch build reports cuda True on the target driver, no cu128 pin needed.
art_pipe stages (concept/modeling/texturing) run on GPU locally and in prod on the updated toolchain.
Follow-up from #425 / spikersoft-backend PR #106 (Windows local ProArt worker).
## Context
When bootstrapping the art_pipe model venvs, art_pipe's venv setup pulled a **CUDA 13.0** torch build (`torch 2.12.1+cu130`), but the dev box's NVIDIA driver only advertises **CUDA 12.9** (driver version 12090). Result: `torch.cuda.is_available()` was `False` and generation would fall back to CPU.
## Current mitigation (shipped in PR #106)
`scripts/windows-artpipe-bootstrap.ps1` pins torch to the **cu128** index (CUDA 12.8 — the newest build compatible with a 12.9 driver, and the index the model manifests already declare). After each model bootstrap it checks CUDA availability in the venv and force-reinstalls torch/torchvision from cu128 if CUDA is unavailable. This keeps GPU generation working today without a host change.
## The CUDA 13 path (this ticket)
We don't want to pin to cu128 forever. Decide + do the migration so the default (cu13x) torch builds work:
- Upgrade the NVIDIA driver on the GPU dev/prod hosts to one that supports CUDA 13.x, then confirm a default `pip install torch` build reports `cuda True`.
- Re-check the CUDA base image tag in `SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile` (currently `nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04`) and align it with the chosen toolkit.
- Confirm the art_pipe model manifests' `torch_index_url` (currently cu128) — bump to cu13x once drivers are updated.
- Once hosts are on a CUDA 13 driver, drop / relax the cu128 pin in `windows-artpipe-bootstrap.ps1`.
- Verify the same holds for the production spikersoft-artpipe-* swarm stacks (SERVER GPU node driver version).
## Acceptance
- Default torch build reports `cuda True` on the target driver, no cu128 pin needed.
- art_pipe stages (concept/modeling/texturing) run on GPU locally and in prod on the updated toolchain.
Board-sweep status (2026-07-22): mitigation shipped (cu128 torch pin in the bootstrap script, backend #106). REMAINING: the actual dev-box CUDA-13 driver migration (ops).
Board-sweep status (2026-07-22): mitigation shipped (cu128 torch pin in the bootstrap script, backend #106). REMAINING: the actual dev-box CUDA-13 driver migration (ops).
This ticket needed changes in more than one repo, so it became one issue per repo:
ArtPipe: spikerj/spikersoft-artpipe#59 — bump the model manifests' env.torch_index_url off cu128 and re-verify the venvs.
Backend: spikerj/spikersoft-backend#608 — align the ArtPipeProcessor CUDA base image and drop the cu128 pin from windows-artpipe-bootstrap.ps1.
The shape of this ticket has changed and both children record it: the driver half is done. nvidia-smi on the GPU host reports NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0 on the RTX 4090. So the ops step this was blocked on has happened, and what remains is purely the code still pinned to the old toolchain.
Verified 2026-08-07: on spikersoft-artpipe@7082d26, grep -rho "cu1[0-9]*" models/*/artpipe.json returns cu128 x 20 — every manifest is still pinned. On spikersoft-backend@98102023, SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile:6 is still nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04 and the ps1 still carries the cu128 pin.
Recorded on both children as a precondition: confirm the SERVER lane also reports CUDA 13.x before bumping — the coordinator now runs two enabled capacity nodes, and #822 left an unverified NVML driver/library mismatch on SERVER. Bumping while SERVER is on 12.x would move the breakage rather than fix it.
Status: partially done — mitigation shipped (backend #106) and the driver upgraded on the 4090; the manifest bump, the base-image bump and the pin removal remain.
Closing here. Work now lives in the repos that hold the fixes, so fixes #N in a PR will auto-close them on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Migrated as part of the umbrella-tracker breakup.
This ticket needed changes in more than one repo, so it became one issue per repo:
- ArtPipe: spikerj/spikersoft-artpipe#59 — bump the model manifests' `env.torch_index_url` off cu128 and re-verify the venvs.
- Backend: spikerj/spikersoft-backend#608 — align the ArtPipeProcessor CUDA base image and drop the cu128 pin from `windows-artpipe-bootstrap.ps1`.
**The shape of this ticket has changed and both children record it: the driver half is done.**
`nvidia-smi` on the GPU host reports `NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0` on the RTX 4090. So the ops step this was blocked on has happened, and what remains is purely the code still pinned to the old toolchain.
Verified 2026-08-07: on `spikersoft-artpipe@7082d26`, `grep -rho "cu1[0-9]*" models/*/artpipe.json` returns cu128 x 20 — every manifest is still pinned. On `spikersoft-backend@98102023`, `SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile:6` is still `nvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04` and the ps1 still carries the cu128 pin.
Recorded on both children as a precondition: confirm the SERVER lane also reports CUDA 13.x before bumping — the coordinator now runs two enabled capacity nodes, and #822 left an unverified NVML driver/library mismatch on SERVER. Bumping while SERVER is on 12.x would move the breakage rather than fix it.
Status: partially done — mitigation shipped (backend #106) and the driver upgraded on the 4090; the manifest bump, the base-image bump and the pin removal remain.
Closing here. Work now lives in the repos that hold the fixes, so `fixes #N` in a PR will auto-close them on merge. The umbrella tracker keeps cross-repo epics only.
— Opus 5 Agent
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Follow-up from #425 / spikersoft-backend PR #106 (Windows local ProArt worker).
Context
When bootstrapping the art_pipe model venvs, art_pipe's venv setup pulled a CUDA 13.0 torch build (
torch 2.12.1+cu130), but the dev box's NVIDIA driver only advertises CUDA 12.9 (driver version 12090). Result:torch.cuda.is_available()wasFalseand generation would fall back to CPU.Current mitigation (shipped in PR #106)
scripts/windows-artpipe-bootstrap.ps1pins torch to the cu128 index (CUDA 12.8 — the newest build compatible with a 12.9 driver, and the index the model manifests already declare). After each model bootstrap it checks CUDA availability in the venv and force-reinstalls torch/torchvision from cu128 if CUDA is unavailable. This keeps GPU generation working today without a host change.The CUDA 13 path (this ticket)
We don't want to pin to cu128 forever. Decide + do the migration so the default (cu13x) torch builds work:
pip install torchbuild reportscuda True.SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile(currentlynvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04) and align it with the chosen toolkit.torch_index_url(currently cu128) — bump to cu13x once drivers are updated.windows-artpipe-bootstrap.ps1.Acceptance
cuda Trueon the target driver, no cu128 pin needed.Board-sweep status (2026-07-22): mitigation shipped (cu128 torch pin in the bootstrap script, backend #106). REMAINING: the actual dev-box CUDA-13 driver migration (ops).
Migrated as part of the umbrella-tracker breakup.
This ticket needed changes in more than one repo, so it became one issue per repo:
env.torch_index_urloff cu128 and re-verify the venvs.windows-artpipe-bootstrap.ps1.The shape of this ticket has changed and both children record it: the driver half is done.
nvidia-smion the GPU host reportsNVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0on the RTX 4090. So the ops step this was blocked on has happened, and what remains is purely the code still pinned to the old toolchain.Verified 2026-08-07: on
spikersoft-artpipe@7082d26,grep -rho "cu1[0-9]*" models/*/artpipe.jsonreturns cu128 x 20 — every manifest is still pinned. Onspikersoft-backend@98102023,SpikerSoft.EventHandlers.ArtPipeProcessor/Dockerfile:6is stillnvidia/cuda:12.9.1-cudnn-devel-ubuntu24.04and the ps1 still carries the cu128 pin.Recorded on both children as a precondition: confirm the SERVER lane also reports CUDA 13.x before bumping — the coordinator now runs two enabled capacity nodes, and #822 left an unverified NVML driver/library mismatch on SERVER. Bumping while SERVER is on 12.x would move the breakage rather than fix it.
Status: partially done — mitigation shipped (backend #106) and the driver upgraded on the 4090; the manifest bump, the base-image bump and the pin removal remain.
Closing here. Work now lives in the repos that hold the fixes, so
fixes #Nin a PR will auto-close them on merge. The umbrella tracker keeps cross-repo epics only.— Opus 5 Agent