CUDA bandwidth king in a small case
Used RTX 5090 SFF workstation
A class, not a single SKU: any small-form-factor box built or bought around a used GeForce RTX 5090. 32GB GDDR7, ~1,792 GB/s, full CUDA stack. The FE is two-slot and SFF-friendly; the power and heat are not.
- Street price
- $2,200–$3,000 card; $3,500–$6,000+ system
- Memory
- 32GB GDDR7
- Bandwidth
- ~1,792 GB/s
- Chip
- Blackwell GB202 (RTX 5090)
- Catch
- 32GB VRAM wall; 575W + heat; used risk
What it is
Any deskside SFF or mini-ITX/ATX build (or used prebuilt) centered on a used NVIDIA GeForce RTX 5090. The card itself is 32GB GDDR7 on a 512-bit bus for ~1,792 GB/s, 21,760 CUDA cores, 575W TGP, PCIe 5.0. Founders Edition is two-slot and the only 5090 that meets Nvidia’s SFF-Ready guidelines; partner cards vary in length and thickness. Pair it with a 1000W+ PSU, decent airflow case, and whatever CPU/RAM/storage fits. This is the discrete-GPU path to local inference, not unified memory.
Who it is for
CUDA people who want maximum tokens-per-second on models that fit in 32GB, and who will live with heat, noise, and the used market. Useful for 8B–32B dense at high speed, image/video workflows, and anything that benefits from full TensorRT / vLLM / CUDA. Wrong if you need 70B+ dense without heavy quant or offload, quiet operation, or a sealed appliance with 128GB unified memory.
What it actually runs
- Comfortable: 8B–32B dense at high token rates (often 100+ tok/s on smaller models); MoE variants that stay inside 32GB.
- Possible: heavier quant 70B-class with some system RAM offload, but decode drops once layers spill over PCIe. Long context eats VRAM fast.
- Not the win: models that need 40GB+ resident VRAM, silent deskside boxes, or low power. Capacity is the hard limit; bandwidth is the strength while the model fits.
Standout details
- Price snapshot early Sep 2026: used 5090 cards tracked around a €2,390 median sold (roughly $2,200–$3,000 USD range depending on market and condition); new street was higher and volatile. Full SFF systems (card + host) commonly land $3,500–$6,000+ once you add case, PSU, CPU, RAM, and storage. Stock and listings move daily.
- 32GB is the ceiling. Dense 70B at usable quant often does not fit fully; offload kills the bandwidth advantage.
- 575W TGP demands a serious PSU and real airflow. True SFF is possible with the FE, but the case will still dump heat into the room.
- Used market risks: no remaining warranty on many cards, potential mining or heavy-use history, scams on marketplace listings. Buy tested with return window when possible.
- Opposite trade-off from DGX Spark / GX10 (128GB, lower bandwidth) and the Strix Halo minis (~256 GB/s, 128GB). The 5090 wins speed on what fits; loses on capacity.
Honest verdict
The used CUDA bandwidth play in a small case. Buy it if you already live in CUDA, need max tokens on 8B–32B class models, and accept 32GB as the hard wall plus heat and used-card risk. If you need 70B+ resident without drama, shop the 128GB unified boxes instead.
Prices and stock move. Affiliate links stay off until this directory has real traffic.