Switch language한국어
Back to the list

A Modder’s RTX 4080 Was Enough to Play AAA Games, But Not to Run LLMs, So He Integrated NVIDIA’s Tesla V100 at a Throwaway Price to Run 27B AI Models

TL;DR AI

Key summary

2 min read
  1. A PC modder paired an RTX 4080 with a used NVIDIA Tesla V100 to boost total usable VRAM for local AI inference.

  2. Using an SXM2-to-PCIe adapter and a modified cooler, the build overcame noise and compatibility issues and reached 32GB of VRAM.

  3. The setup was able to run a quantized Qwen3.6-27B model at usable token rates.

  4. The entire upgrade reportedly cost under $300 in parts, showing a low-cost way to repurpose old enterprise GPUs for LLMs.

Read the original