A Modder’s RTX 4080 Was Enough to Play AAA Games, But Not to Run LLMs, So He Integrated NVIDIA’s Tesla V100 at a Throwaway Price to Run 27B AI Models

TL;DR AI
2 min readKey summary
A PC modder paired an RTX 4080 with a used NVIDIA Tesla V100 to boost total usable VRAM for local AI inference.
Using an SXM2-to-PCIe adapter and a modified cooler, the build overcame noise and compatibility issues and reached 32GB of VRAM.
The setup was able to run a quantized Qwen3.6-27B model at usable token rates.
The entire upgrade reportedly cost under $300 in parts, showing a low-cost way to repurpose old enterprise GPUs for LLMs.



