Original post on X
This rig was shared first on X. Reply there or discuss it below.
View original post
About this rig
Shared first on X by Peasant Smith (@PeasantSmith). A single old-generation Tesla V100 32GB paired with a Xeon E5-2690 v4 and 16GB of DDR4 RAM. Loaded a 75.2GB Qwen3.8-Flash-Next quant that doesn’t fully fit in VRAM and still got 24.9 tok/s decoding via CPU/GPU offload - a good example of squeezing a large model onto old enterprise hardware.
Original post: https://x.com/PeasantSmith/status/2097988456815157738
Claim your /u/name
Then build and share a rig like this one - free