docs/3090-SIM.md

3090 as a T4000 sim

Local brick: fractal1, Arch, ssh fractal1 (user dorje). RTX 3090 24 GB, 16 GB host RAM. Not the quality box. Not a T4000. Not a 30 fps host.

The live mix is designed everything resident at 30 fps. This box emulates that pipeline by loading one SMALL slot at a time and unloading it. Slow on purpose. Same software as a local T4000 or a rented RTX PRO 6000 Blackwell Server 96 GB, which keep slots resident.

A PRO 6000 stream is a LARGE-pack lab result. There is no derate from 6000 fps to T4000 (1792 vs 273 GB/s). T4000 30 fps stays SPEC until measured on Thor silicon. No box in the four-envelope list validates resident-set fps at 273 GB/s except Thor itself.

Swap-mode logs SHALL split load/unload time from per-frame time before ×0.13 / ×0.30 apply. Do not derate a number that still includes weight load (the klein row below currently conflates; next measurement splits it).

Cloud T4000-shaped runs stay on RunPod (Blackwell, FP8/FP4). This box is HEVC plumbing + FP16 nets + the translation table below.

T4000 FACT is DS-11945-001 v1.4 @ 70 W. See references/T4000.md. Product sim: web3d-space/docs/all-systems-go/GPU-SIM.md.

Do not divide by CUDA cores

RTX 3090T4000 @ 70 W3090 / T4000
ArchAmpere sm_86, 3rd-gen tensorBlackwell sm_110, 5th-gen + TMEMdifferent
CUDA cores104961536 (2 GPC / 6 TPC)6.8× — wrong axis
FP32 CUDA35.6 TFLOPS4.7 TFLOPS7.6×
FP16 tensor dense / sparse~142 / ~285 TFLOPS~150 / ~300 TFLOPS~1×
INT8 (= FP8 on Thor) dense / sparse~285 / ~570 TOPS300 / 600 TOPS~1×
FP8 sparseno hardware600 TFLOPST4000 only
FP4 sparseno hardware1200 TFLOPST4000 only
Bandwidth936 GB/s GDDR6X273 GB/s LPDDR5X3090 3.4×
Capacity24 GB discrete + 16 GB host64 GB unifiedT4000 wins envelope
NVENC1× Ampere1× Blackwell, HQ 2× 4Kp30same count; 3090 overstates headroom
PVA / OFAnonePVA 3.0 + optical flownot emulatable
Power350 W70 W / 90 W throttledo not nvidia-smi -pl 70 and call it T4000

Datasheet: FP8 TFLOPs is equivalent to INT8 TOPs.

Translation table (read 3090 numbers as T4000)

If you measured on 3090…Multiply byWhy
CUDA-core filter fps (color, warp, undistort, RVM conv)× 0.13FP32 4.7 / 35.6
SAM2 / ViT / LLM prefill at FP16× ~0.30bandwidth 273 / 936; tensor ~1×
LLM decode tok/s× ~0.30decode is GB/s
FP8 / NVFP4 TRTcannotThor-only. 3090 understates production Thor
Fits in 24 GBfits T4000 64 GB3090 is the pessimistic envelope
OOM on 24 GBunknown on T40009B + klein + SAM2 is a 64 GB yes, 24 GB maybe-not
3× 4K NVENCfail T4000 HQT4000 HQ is 2× 4Kp30

Do not derate by 6.8 (core count). Do not treat 16 GB host RAM as a T4000 limit — Thor unified 64 GB eats the “host” half.

Live stack we care about (NVENC + SAM2-tiny + DA-S + one of 9B / klein) is a 273 GB/s and 70 W problem on Thor. 3090 makes the nets look cheap and the memory envelope look expensive. Trust the second, haircut the first.

Ideal next hardware (RunPod)

No RunPod SKU is 6 TPC / 273 GB/s / 70 W. Pick by which lie you are correcting. This 3090 cannot do FP8/FP4; that is the cloud job.

JobRentWhyNot
Quant / kernels (FP8 SAM2, NVFP4 klein, TRT)RTX PRO 4000 Blackwell 24 GB (~$0.57/hr) or keep PRO 4500 32 GB (lduog58vatxh44, $0.72/hr, EU-RO-1)Same 5th-gen tensor as T4000. 4000 is the cheapest Blackwell. 4500 is already the /gpu worker.Tesla T4, Ampere 3090/A6000, Ada L40S as a T4000
30 fps all-filters resident (LARGE pack)RTX PRO 6000 Blackwell Server 96 GB (~$2.09/hr)The lab SKU. Same software as T4000. Capacity holds. TOPS will lie highno derate to T4000.24/32 GB cards for the co-resident test
Memory envelope (9B + klein + SAM2 co-resident, the 64 GB question)same PRO 6000 96 GBOnly common discrete that holds ≥64 GB.
Quality / DiTtruck PRO 6000 — already the quality box, not a T4000 sim

Next dedicated T4000-shaped cloud run: PRO 4000 Blackwell if we only need FP8/FP4 kernels; PRO 6000 Blackwell if we are testing the resident LARGE pack. Do not terminate negotiated-gpu-drafttrain-*. Auto-off 30 min on aicam-* pods. Do not rent until someone says go.

Tesla T4 is Turing 16 GB. Wrong generation, wrong memory, wrong encode. We are not renting one.

Local harness

sim/3090/ on this repo, synced to dorje@fractal1:~/aicam/. Stages unload between jobs (16 GB host).

ssh fractal1
cd ~/aicam && ./run.sh          # filters → SAM2 → CLIP regions → klein snap

Apply the table to any fps / tok/s we print. Fill Measured when a stage lands.

Measured

fractal1, 2026-09-09. SAM2 bedroom.mp4, 48 frames @ 960×540. Harness ~/aicam/.

Stage3090 (FACT)T4000 read (table)Notes
CUDA-core look-dev (unsharp, grade, vignette, grain)266–281 fps, ~700 MB, 44–59 W×0.13 → ~35 fpsLIVE on T4000. Easy.
SAM2.1 Hiera-tiny AMG (float32)1.46–1.53 s/frame (~0.67 fps), ~2.0–3.0 GB, ~390–407 W×0.30 → ~0.20 fpsThis is AMG, not live track. Tiny only returned 2 corner masks on this clip (lamp, floor). Live path is detector-keyframe + tiny tracker.
CLIP ViT-B/32 region embed4–7 ms for 2 cropsbandwidth-bound, still cheapMaps mask crop → prototype (a lamp, a floor). Cosine ~0.23–0.25 — weak because AMG missed the subjects.
FLUX.2 klein 4B 768², 4 steps, cpu_offloadt2i 43 s (load + warmup in that); i2i 36 sFP16 ~1×; Thor FP8/NVFP4 may beatApache 4B. 16 GB host survived via offload. Unload SAM2 first.
Qwen3.5-9Bnot this passdecode ×0.30Host RAM. Next: after klein unload, or RunPod 96 GB envelope.

Bay glass

The live bay is its own repo now: fractal1-web (ssh://git@mimir.worldtree.network/duke/fractal1-web.git, on fractal1 at ~/work/fractal1-web). It owns the feed (feed/serve.py, unit fractal1-web-feed), nginx, the build timer, and the public tunnel: read-only https://fractal1.vm.worldtree.network, permissive http://fractal1/ on the tailnet. See that repo’s AGENTS.md.

In this repo, vite dev still proxies /bay-api to fractal1:8745; GitHub Pages and other static builds show the last plate.

Refresh the plate:

ssh fractal1 python3 ~/work/fractal1-web/feed/snapshot.py > src/lib/bay/last-plate.json

Derates ×0.13 / ×0.30 live in this file and src/lib/bay/t4000.ts. The feed does not emit them.