Pixal3D PR-5: deploy stack, method option, and routing cutover #838

Closed
opened 2026-07-25 00:54:08 +00:00 by spikerj · 4 comments
Owner

Part of #833. Blocked by #836 (routing allow-list) and #837 (env image).

Stack

spikersoft-infrastructure/spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml, copied from spikersoft-artpipe-model-triposr:

  • Pin node.labels.artpipe-gpu == true; keep the env-var GPU access pattern (NVIDIA_VISIBLE_DEVICES=all — compose v3 rejects deploy.resources.reservations.devices).
  • ArtPipe__ModelQueue=Pixal3D, ArtPipe__Stages=
  • ArtPipe__StageSettings__modeling__{ModelDir=Pixal3D,Action=image_to_3d,VramMB=18000,TimeoutSeconds=7200}
  • Storage__UseS3=true, HF_HUB_OFFLINE=1
  • Join the jaeger and seq-attachable overlays — exporter config without network membership silently drops traces.

Method option

Add to ArtStudioStageMethodOptions.CreateDefault()'s Modeling list:

new() { Key = "pixal3d", ModelDir = "Pixal3D", Action = "image_to_3d", Label = "Pixal3D (highest fidelity)" }

Not first — the first entry is the default and must mirror the worker's StageSettings. No AppleSilicon variant (unlike the trellis key). Edit the C# default, not the ArtStudio:StageMethodOptions config section — Bind replaces defaults wholesale, so a config-only entry would wipe every other stage's methods.

Do not use "disabled": true as the pre-launch gate: it's honored by :9100 but ignored by IsRunnableModelDir. Gate by simply not adding the method option until the image ships.

Verification

  • Submit one TextPrompt and one ImageUpload asset with pixal3d selected — they take different publish paths, so testing one proves nothing about the other
  • For each: single Jaeger trace spanning API → RabbitMQ → worker (proves traceparent survived); GLB in the art-asset-artifacts MinIO bucket; UI advancing via ReceiveArtAssetStageProgress; art.asset.unrouted.messages empty
  • Regression: a TripoSG-selected modeling job still runs on spikersoft-artpipe-modeling
  • Confirm the 18,000 MB lease is granted on the lane, and note the serialization with spikersoft-image-description (20,000 MB on the same card)
Part of #833. Blocked by #836 (routing allow-list) and #837 (env image). ## Stack `spikersoft-infrastructure/spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml`, copied from `spikersoft-artpipe-model-triposr`: - Pin `node.labels.artpipe-gpu == true`; keep the env-var GPU access pattern (`NVIDIA_VISIBLE_DEVICES=all` — compose v3 rejects `deploy.resources.reservations.devices`). - `ArtPipe__ModelQueue=Pixal3D`, `ArtPipe__Stages=` - `ArtPipe__StageSettings__modeling__{ModelDir=Pixal3D,Action=image_to_3d,VramMB=18000,TimeoutSeconds=7200}` - `Storage__UseS3=true`, `HF_HUB_OFFLINE=1` - Join the `jaeger` and `seq-attachable` overlays — **exporter config without network membership silently drops traces.** ## Method option Add to `ArtStudioStageMethodOptions.CreateDefault()`'s Modeling list: ```csharp new() { Key = "pixal3d", ModelDir = "Pixal3D", Action = "image_to_3d", Label = "Pixal3D (highest fidelity)" } ``` **Not first** — the first entry is the default and must mirror the worker's StageSettings. **No `AppleSilicon` variant** (unlike the `trellis` key). Edit the C# default, not the `ArtStudio:StageMethodOptions` config section — `Bind` replaces defaults wholesale, so a config-only entry would wipe every other stage's methods. Do **not** use `"disabled": true` as the pre-launch gate: it's honored by `:9100` but ignored by `IsRunnableModelDir`. Gate by simply not adding the method option until the image ships. ## Verification - [ ] Submit one **TextPrompt** and one **ImageUpload** asset with `pixal3d` selected — they take different publish paths, so testing one proves nothing about the other - [ ] For each: single Jaeger trace spanning API → RabbitMQ → worker (proves `traceparent` survived); GLB in the `art-asset-artifacts` MinIO bucket; UI advancing via `ReceiveArtAssetStageProgress`; `art.asset.unrouted.messages` empty - [ ] Regression: a TripoSG-selected modeling job still runs on `spikersoft-artpipe-modeling` - [ ] Confirm the 18,000 MB lease is granted on the lane, and note the serialization with `spikersoft-image-description` (20,000 MB on the same card)
Author
Owner

Scope shrank a lot — #836 turned out to be already implemented on master under #357, so there is no backend code change needed for routing. Pixal3D is just another entry in the established per-model modeling lane (triposg / shape / instantmesh / sf3d / hunyuan).

Stack file landed in spikersoft-infrastructure PR #149 — a copy of spikersoft-artpipe-model-triposg with ArtPipe__ModelQueue=Pixal3D and VramMB=18000. It is inert until the deploy loop references it.

Corrected remaining checklist

Sequenced so nothing deploys before it's proven:

  1. Merge spikersoft-artpipe #33 (manifest + backend + env-image entry) and #34 (VRAM gate).
  2. workflow_dispatch → ModelEnvImages with images: "pixal3d". First real test of the setup chain — TRELLIS.2 extension builds, natten from source, the ~35-45 GB bake. install.verify_imports should fail the job rather than ship a green image if any nvcc build failed.
  3. workflow_dispatch → ArtPipeProcessor with images: "pixal3d" for the tier-3 .NET layer.
  4. Add spikersoft-artpipe-model-pixal3d to the deploy loop in .gitea/workflows/spikersoft-artpipe-processor.yml (one line, alongside the other per-model stacks).
  5. Only then add the method option to ArtStudioStageMethodOptions.CreateDefault()'s Modeling list — not first in the list, no AppleSilicon variant. Gating by "don't add the option until the image ships" rather than by "disabled": true, which :9100 honors but IsRunnableModelDir ignores.
  6. Routing: set modeling in ArtStudio:StageModelMap (API, first-stage/restart publishes) and ModelQueueStages on the publishing stack (worker, next-stage hop). Both already exist — but note they route the whole modeling stage to per-model queues, so every other modeling model must already have its container up. It does today.

Deliberately not in DEFAULT_IMAGES on either workflow: at ~35-45 GB it's the heaviest image in the lane, so it stays build-on-demand.

**Scope shrank a lot** — #836 turned out to be already implemented on `master` under #357, so there is **no backend code change** needed for routing. Pixal3D is just another entry in the established per-model modeling lane (`triposg` / `shape` / `instantmesh` / `sf3d` / `hunyuan`). Stack file landed in spikersoft-infrastructure PR #149 — a copy of `spikersoft-artpipe-model-triposg` with `ArtPipe__ModelQueue=Pixal3D` and `VramMB=18000`. It is inert until the deploy loop references it. ## Corrected remaining checklist Sequenced so nothing deploys before it's proven: 1. Merge spikersoft-artpipe #33 (manifest + backend + env-image entry) and #34 (VRAM gate). 2. `workflow_dispatch` → ModelEnvImages with `images: "pixal3d"`. **First real test of the setup chain** — TRELLIS.2 extension builds, natten from source, the ~35-45 GB bake. `install.verify_imports` should fail the job rather than ship a green image if any nvcc build failed. 3. `workflow_dispatch` → ArtPipeProcessor with `images: "pixal3d"` for the tier-3 .NET layer. 4. Add `spikersoft-artpipe-model-pixal3d` to the deploy loop in `.gitea/workflows/spikersoft-artpipe-processor.yml` (one line, alongside the other per-model stacks). 5. **Only then** add the method option to `ArtStudioStageMethodOptions.CreateDefault()`'s Modeling list — not first in the list, no `AppleSilicon` variant. Gating by "don't add the option until the image ships" rather than by `"disabled": true`, which `:9100` honors but `IsRunnableModelDir` ignores. 6. Routing: set `modeling` in `ArtStudio:StageModelMap` (API, first-stage/restart publishes) and `ModelQueueStages` on the publishing stack (worker, next-stage hop). Both already exist — but note they route the *whole modeling stage* to per-model queues, so every other modeling model must already have its container up. It does today. Deliberately **not** in `DEFAULT_IMAGES` on either workflow: at ~35-45 GB it's the heaviest image in the lane, so it stays build-on-demand.
Author
Owner

Audited against origin/master / origin/main — PARTIAL, and the infra stack is currently inert. Staying open. Concrete remaining items:

Landed

  • spikersoft-infrastructure spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml (infra PR #149, 3921bd8) — ArtPipe__ModelQueue=Pixal3D, empty ArtPipe__Stages, VramMB=18000, artpipe-gpu node label, env-var GPU access, HF_HUB_OFFLINE=1.
  • Tier-2 build case: spikersoft-artpipe/.gitea/workflows/model-env-images.yml:269 — pixal3d) MODELS="Pixal3D" ;;, deliberately out of DEFAULT_IMAGES (:220) as the comment specifies.
  • Routing needed no backend change — #836 was already implemented under #357. Correct per the earlier comment.

Missing. git grep -in "pixal" origin/master returns zero matches in spikersoft-backend (and zero in spikersoft-angular):

  1. Tier-3 image case — no pixal3d) branch in .gitea/workflows/spikersoft-artpipe-processor.yml; its DEFAULT_IMAGES (:188) lists ten images with no pixal3d, and there is no case arm to build it even on manual dispatch.
  2. Deploy-loop entry — .gitea/workflows/spikersoft-artpipe-processor.yml:548-553 enumerates the per-model stacks and spikersoft-artpipe-model-pixal3d is absent. This is why the merged stack file does nothing today — exactly as the earlier comment predicted.
  3. Method option — no Key = "pixal3d" entry in ArtStudioStageMethodOptions.CreateDefault()'s Modeling list, so it can't be selected.

All verification checkboxes (TextPrompt + ImageUpload runs, Jaeger trace, MinIO GLB, lease grant) remain unrun — unsurprising, since nothing can currently route to it.

Note for sequencing: #834 and #835 are code-complete but their own 4090 hardware runs are also unrun, so this PR-5 cutover is the gate for validating all three.

Audited against `origin/master` / `origin/main` — **PARTIAL, and the infra stack is currently inert.** Staying open. Concrete remaining items: **Landed** - `spikersoft-infrastructure` `spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml` (infra PR #149, `3921bd8`) — `ArtPipe__ModelQueue=Pixal3D`, empty `ArtPipe__Stages`, `VramMB=18000`, `artpipe-gpu` node label, env-var GPU access, `HF_HUB_OFFLINE=1`. - Tier-2 build case: `spikersoft-artpipe/.gitea/workflows/model-env-images.yml:269` — `pixal3d) MODELS="Pixal3D" ;;`, deliberately out of `DEFAULT_IMAGES` (`:220`) as the comment specifies. - Routing needed no backend change — #836 was already implemented under #357. Correct per the earlier comment. **Missing.** `git grep -in "pixal" origin/master` returns **zero matches in spikersoft-backend** (and zero in spikersoft-angular): 1. **Tier-3 image case** — no `pixal3d)` branch in `.gitea/workflows/spikersoft-artpipe-processor.yml`; its `DEFAULT_IMAGES` (`:188`) lists ten images with no pixal3d, and there is no case arm to build it *even on manual dispatch*. 2. **Deploy-loop entry** — `.gitea/workflows/spikersoft-artpipe-processor.yml:548-553` enumerates the per-model stacks and `spikersoft-artpipe-model-pixal3d` is absent. **This is why the merged stack file does nothing today** — exactly as the earlier comment predicted. 3. **Method option** — no `Key = "pixal3d"` entry in `ArtStudioStageMethodOptions.CreateDefault()`'s Modeling list, so it can't be selected. All verification checkboxes (TextPrompt + ImageUpload runs, Jaeger trace, MinIO GLB, lease grant) remain unrun — unsurprising, since nothing can currently route to it. Note for sequencing: #834 and #835 are code-complete but their own 4090 hardware runs are also unrun, so this PR-5 cutover is the gate for validating all three.
Author
Owner

Code is done, deploy and cutover are not.

I closed #834 and #835 — the artpipe manifest/runner/backend and the tunable VRAM gate are on master, and the stack file spikersoft-infrastructure/spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml exists (infra#149).

Two things in this ticket are still outstanding, verified just now:

  1. Not deployed. docker service ls shows the artpipe model lanes as blender, florence2, hunyuan, instantmesh, qrmonster, safety, sdxl, sf3d, shape, textto3d, triposg, triposr, photostack — no spikersoft-artpipe-model-pixal3d.
  2. No method option. grep -ri pixal across spikersoft-backend returns nothing, so Pixal3D was never added to ArtStudioStageMethodOptions.CreateDefault()'s Modeling list and can't be selected.

Epic #833 stays blocked on this.

— 2026-08-06 tracker sweep, Opus 5 Agent. Staying open.

**Code is done, deploy and cutover are not.** I closed #834 and #835 — the artpipe manifest/runner/backend and the tunable VRAM gate are on master, and the stack file `spikersoft-infrastructure/spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml` exists (infra#149). Two things in this ticket are still outstanding, verified just now: 1. **Not deployed.** `docker service ls` shows the artpipe model lanes as blender, florence2, hunyuan, instantmesh, qrmonster, safety, sdxl, sf3d, shape, textto3d, triposg, triposr, photostack — **no `spikersoft-artpipe-model-pixal3d`**. 2. **No method option.** `grep -ri pixal` across spikersoft-backend returns nothing, so Pixal3D was never added to `ArtStudioStageMethodOptions.CreateDefault()`'s Modeling list and can't be selected. Epic #833 stays blocked on this. — 2026-08-06 tracker sweep, Opus 5 Agent. Staying open.
Author
Owner

Migrated to spikerj/spikersoft-backend#570 as part of the umbrella-tracker breakup.

This did not need a split — the infra half (spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml, infra PR #149 3921bd8) and the artpipe half (models/Pixal3D/artpipe.json + the pixal3d) case at model-env-images.yml:269) are already on their default branches, and routing needed no backend change since #836 was already implemented under #357. Everything outstanding lives in spikersoft-backend.

Verified 2026-08-07 against spikersoft-backend@98102023: grep -rin "pixal" returns zero matches. No pixal3d) tier-3 build case, no spikersoft-artpipe-model-pixal3d in the deploy loop at .gitea/workflows/spikersoft-artpipe-processor.yml:552-558, and no Key = "pixal3d" in ArtStudioStageMethodOptions.CreateDefault() (ArtStudioStageMethodOptions.cs:89).

Live: docker service ls shows thirteen artpipe lanes (blender, florence2, hunyuan, instantmesh, qrmonster, safety, sdxl, sf3d, shape, textto3d, triposg, triposr, photostack) plus artpipe-modeling — no pixal3d service exists, so the merged stack file is inert and every verification checkbox is unrun.

Status: partially done — infra + artpipe merged; all three backend items and the whole verification pass still missing.

Closing here. Work now lives in the repo that holds the fix, so fixes #570 in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only.

— Opus 5 Agent

Migrated to **spikerj/spikersoft-backend#570** as part of the umbrella-tracker breakup. This did **not** need a split — the infra half (`spikersoft-artpipe-model-pixal3d/docker-stack-gpu.yml`, infra PR #149 `3921bd8`) and the artpipe half (`models/Pixal3D/artpipe.json` + the `pixal3d)` case at `model-env-images.yml:269`) are already on their default branches, and routing needed no backend change since #836 was already implemented under #357. Everything outstanding lives in spikersoft-backend. Verified 2026-08-07 against `spikersoft-backend@98102023`: `grep -rin "pixal"` returns zero matches. No `pixal3d)` tier-3 build case, no `spikersoft-artpipe-model-pixal3d` in the deploy loop at `.gitea/workflows/spikersoft-artpipe-processor.yml:552-558`, and no `Key = "pixal3d"` in `ArtStudioStageMethodOptions.CreateDefault()` (`ArtStudioStageMethodOptions.cs:89`). Live: `docker service ls` shows thirteen artpipe lanes (blender, florence2, hunyuan, instantmesh, qrmonster, safety, sdxl, sf3d, shape, textto3d, triposg, triposr, photostack) plus `artpipe-modeling` — **no pixal3d service exists**, so the merged stack file is inert and every verification checkbox is unrun. Status: partially done — infra + artpipe merged; all three backend items and the whole verification pass still missing. Closing here. Work now lives in the repo that holds the fix, so `fixes #570` in a PR will auto-close it on merge. The umbrella tracker keeps cross-repo epics only. — Opus 5 Agent
Sign in to join this conversation.