docs/3090-SIM.md
3090 as a T4000 sim
Local brick: fractal1, Arch, ssh fractal1 (user dorje). RTX 3090 24 GB, 16 GB host RAM. Not the quality box. Not a T4000. Not a 30 fps host.
The live mix is designed everything resident at 30 fps. This box emulates that pipeline by loading one SMALL slot at a time and unloading it. Slow on purpose. Same software as a local T4000 or a rented RTX PRO 6000 Blackwell Server 96 GB, which keep slots resident.
A PRO 6000 stream is a LARGE-pack lab result. There is no derate from 6000 fps to T4000 (1792 vs 273 GB/s). T4000 30 fps stays SPEC until measured on Thor silicon. No box in the four-envelope list validates resident-set fps at 273 GB/s except Thor itself.
Swap-mode logs SHALL split load/unload time from per-frame time before ×0.13 / ×0.30 apply. Do not derate a number that still includes weight load (the klein row below currently conflates; next measurement splits it).
Cloud T4000-shaped runs stay on RunPod (Blackwell, FP8/FP4). This box is HEVC plumbing + FP16 nets + the translation table below.
T4000 FACT is DS-11945-001 v1.4 @ 70 W. See references/T4000.md. Product sim: web3d-space/docs/all-systems-go/GPU-SIM.md.
Do not divide by CUDA cores
| RTX 3090 | T4000 @ 70 W | 3090 / T4000 | |
|---|---|---|---|
| Arch | Ampere sm_86, 3rd-gen tensor | Blackwell sm_110, 5th-gen + TMEM | different |
| CUDA cores | 10496 | 1536 (2 GPC / 6 TPC) | 6.8× — wrong axis |
| FP32 CUDA | 35.6 TFLOPS | 4.7 TFLOPS | 7.6× |
| FP16 tensor dense / sparse | ~142 / ~285 TFLOPS | ~150 / ~300 TFLOPS | ~1× |
| INT8 (= FP8 on Thor) dense / sparse | ~285 / ~570 TOPS | 300 / 600 TOPS | ~1× |
| FP8 sparse | no hardware | 600 TFLOPS | T4000 only |
| FP4 sparse | no hardware | 1200 TFLOPS | T4000 only |
| Bandwidth | 936 GB/s GDDR6X | 273 GB/s LPDDR5X | 3090 3.4× |
| Capacity | 24 GB discrete + 16 GB host | 64 GB unified | T4000 wins envelope |
| NVENC | 1× Ampere | 1× Blackwell, HQ 2× 4Kp30 | same count; 3090 overstates headroom |
| PVA / OFA | none | PVA 3.0 + optical flow | not emulatable |
| Power | 350 W | 70 W / 90 W throttle | do not nvidia-smi -pl 70 and call it T4000 |
Datasheet: FP8 TFLOPs is equivalent to INT8 TOPs.
Translation table (read 3090 numbers as T4000)
| If you measured on 3090… | Multiply by | Why |
|---|---|---|
| CUDA-core filter fps (color, warp, undistort, RVM conv) | × 0.13 | FP32 4.7 / 35.6 |
| SAM2 / ViT / LLM prefill at FP16 | × ~0.30 | bandwidth 273 / 936; tensor ~1× |
| LLM decode tok/s | × ~0.30 | decode is GB/s |
| FP8 / NVFP4 TRT | cannot | Thor-only. 3090 understates production Thor |
| Fits in 24 GB | fits T4000 64 GB | 3090 is the pessimistic envelope |
| OOM on 24 GB | unknown on T4000 | 9B + klein + SAM2 is a 64 GB yes, 24 GB maybe-not |
| 3× 4K NVENC | fail T4000 HQ | T4000 HQ is 2× 4Kp30 |
Do not derate by 6.8 (core count). Do not treat 16 GB host RAM as a T4000 limit — Thor unified 64 GB eats the “host” half.
Live stack we care about (NVENC + SAM2-tiny + DA-S + one of 9B / klein) is a 273 GB/s and 70 W problem on Thor. 3090 makes the nets look cheap and the memory envelope look expensive. Trust the second, haircut the first.
Ideal next hardware (RunPod)
No RunPod SKU is 6 TPC / 273 GB/s / 70 W. Pick by which lie you are correcting. This 3090 cannot do FP8/FP4; that is the cloud job.
| Job | Rent | Why | Not |
|---|---|---|---|
| Quant / kernels (FP8 SAM2, NVFP4 klein, TRT) | RTX PRO 4000 Blackwell 24 GB (~$0.57/hr) or keep PRO 4500 32 GB (lduog58vatxh44, $0.72/hr, EU-RO-1) | Same 5th-gen tensor as T4000. 4000 is the cheapest Blackwell. 4500 is already the /gpu worker. | Tesla T4, Ampere 3090/A6000, Ada L40S as a T4000 |
| 30 fps all-filters resident (LARGE pack) | RTX PRO 6000 Blackwell Server 96 GB (~$2.09/hr) | The lab SKU. Same software as T4000. Capacity holds. TOPS will lie high — no derate to T4000. | 24/32 GB cards for the co-resident test |
| Memory envelope (9B + klein + SAM2 co-resident, the 64 GB question) | same PRO 6000 96 GB | Only common discrete that holds ≥64 GB. | — |
| Quality / DiT | truck PRO 6000 — already the quality box, not a T4000 sim | — | — |
Next dedicated T4000-shaped cloud run: PRO 4000 Blackwell if we only need FP8/FP4 kernels; PRO 6000 Blackwell if we are testing the resident LARGE pack. Do not terminate negotiated-gpu-drafttrain-*. Auto-off 30 min on aicam-* pods. Do not rent until someone says go.
Tesla T4 is Turing 16 GB. Wrong generation, wrong memory, wrong encode. We are not renting one.
Local harness
sim/3090/ on this repo, synced to dorje@fractal1:~/aicam/. Stages unload between jobs (16 GB host).
ssh fractal1
cd ~/aicam && ./run.sh # filters → SAM2 → CLIP regions → klein snap Apply the table to any fps / tok/s we print. Fill Measured when a stage lands.
Measured
fractal1, 2026-09-09. SAM2 bedroom.mp4, 48 frames @ 960×540. Harness ~/aicam/.
| Stage | 3090 (FACT) | T4000 read (table) | Notes |
|---|---|---|---|
| CUDA-core look-dev (unsharp, grade, vignette, grain) | 266–281 fps, ~700 MB, 44–59 W | ×0.13 → ~35 fps | LIVE on T4000. Easy. |
| SAM2.1 Hiera-tiny AMG (float32) | 1.46–1.53 s/frame (~0.67 fps), ~2.0–3.0 GB, ~390–407 W | ×0.30 → ~0.20 fps | This is AMG, not live track. Tiny only returned 2 corner masks on this clip (lamp, floor). Live path is detector-keyframe + tiny tracker. |
| CLIP ViT-B/32 region embed | 4–7 ms for 2 crops | bandwidth-bound, still cheap | Maps mask crop → prototype (a lamp, a floor). Cosine ~0.23–0.25 — weak because AMG missed the subjects. |
| FLUX.2 klein 4B 768², 4 steps, cpu_offload | t2i 43 s (load + warmup in that); i2i 36 s | FP16 ~1×; Thor FP8/NVFP4 may beat | Apache 4B. 16 GB host survived via offload. Unload SAM2 first. |
| Qwen3.5-9B | not this pass | decode ×0.30 | Host RAM. Next: after klein unload, or RunPod 96 GB envelope. |
Bay glass
The live bay is its own repo now: fractal1-web (ssh://git@mimir.worldtree.network/duke/fractal1-web.git, on fractal1 at ~/work/fractal1-web). It owns the feed (feed/serve.py, unit fractal1-web-feed), nginx, the build timer, and the public tunnel: read-only https://fractal1.vm.worldtree.network, permissive http://fractal1/ on the tailnet. See that repo’s AGENTS.md.
In this repo, vite dev still proxies /bay-api to fractal1:8745; GitHub Pages and other static builds show the last plate.
Refresh the plate:
ssh fractal1 python3 ~/work/fractal1-web/feed/snapshot.py > src/lib/bay/last-plate.json Derates ×0.13 / ×0.30 live in this file and src/lib/bay/t4000.ts. The feed does not emit them.