A Tesla V100 purchased for around US$200 became a video memory expansion in a gaming PC that was already running AAA games with an RTX 4080. The set included 32 GB of VRAM and started running a language model with 27 billion parameters at 32 tokens per second, without depending on the cloud or internet connection.
The editing is by Oscar Molnar, who documented the process in your blog however, he emphasizes that GeForce is still the one who designs the frames. The server board came in only to load the model into memory.
What cost $266
The total disbursement was £200, something close to US$266. The approximate division looks like this, with conversion at the commercial rate of R$5.08 this Monday (27):
| Item | Dollar value | Approximate conversion |
|---|---|---|
| Tesla V100 SXM2 16GB | Around US$200 | R$ 1,016 |
| SXM2 to PCIe Adapter | About US$66 | R$ 335 |
| Total | $266 | R$ 1,351 |
The board does not have a PCIe slot, video output or power connector
The Tesla V100 used is the SXM2 variant, a format that NVIDIA created for modules slotted directly into server baseboards. It has no PCIe contact edge, no PCIe power connector and no video output.
Installing this card in a home machine requires an SXM2 to PCIe adapter, a part sold on the black market for around US$66. It is the item that translates the server socket into the x16 slot on a common motherboard.
There is a 32 GB version of the same card, which would double the memory in one go. The price, however, also doubles, and it circulates in the US$400 to US$500 range on the second-hand market.
82 dB and the fan mod
The original V100 SXM2 cooler was designed for racks, where no one can hear anything. Measured in a home environment, it recorded 82 dB. Molnar described the result as something “between a garbage disposal and a lawnmower.”

The solution was simple and did not require an expensive part, the fan wires were redirected to the motherboard’s PWM header, which allowed the operating system to control the rotation. It also works to buy a 2.54 mm male to PH2.0 female jumper cable and bridge it.
With rotation control available, the fan started to rotate at 10% of maximum speed. Even at this level, the card remains below 50 °C under full load, which provides a comfortable margin for prolonged operation.
How Volta and Ada live on the same machine
The other obstacle was software: the PC combines two architectures separated by five years: the RTX 4080 is Ada Lovelace, launched in 2022, and the Tesla V100 is Volta, from 2017.
The solution was to run NixOS with a legacy NVIDIA driver, a version whose compatibility window still covers both architectures at the same time. From there, the system sees a combined 32 GB of VRAM and distributes the model layers between the two cards for local inference.
The V100 continues to deliver respectable numbers almost a decade after launch. There are 5,120 CUDA cores and a 4,096-bit bus that supports around 900 GB/s of bandwidth, a level that most consumer GeForce devices do not reach.
| Specification | Tesla V100 SXM2 |
|---|---|
| Architecture | Volta (GV100) |
| CUDA Cores | 5,120 |
| Memory | 16GB HBM2 |
| Memory bus | 4,096 bits |
| Bandwidth | Around 900 GB/s |
| Tensor Cores | First generation |
| Ray Tracing Cores | None |
| Video output | None |
| Physical interface | SXM2, requires adapter |
The absence of cores dedicated to Ray Tracing and any video output explains why the card does not replace a GeForce.
In language model inference, however, the math changes: the task is limited by memory bandwidth, and that is precisely where HBM2 has an advantage over GDDR6 on newer mid-range cards.

Qwen3.6-27B at 32 tokens per second
The test ran Qwen3.6-27B-MTP quantized in Q5_K_M, 19 GB file, with a context window of 128 thousand tokens. There was memory left for the entire context without having to use the system’s RAM.
| Metric | Result |
|---|---|
| Model | Qwen3.6-27B-MTP |
| Quantization | Q5_K_M |
| File size | 19GB |
| Context window | 128 thousand tokens |
| Text generation | 32 tokens per second |
| Prompt processing | 133 to 160 tokens per second |
| Total system VRAM | 32 GB (16 GB from RTX 4080 plus 16 GB from V100) |
“Fast enough for interactive use”, summarized Oscar Molnar in the report published on his blog, adding that the rate obtained exceeds that of most alternatives accessed via API in the cloud.
For comparison, comfortable human reading is around 10 to 15 tokens per second. The 32 tokens per second puts the response above the reading speed, which eliminates the feeling of waiting while typing output.
Also read:
- 8-year-old NVIDIA V100 GPU becomes AI phenomenon after dropping to $100
- Micron launches 3GB GDDR7 memory to increase VRAM in entry-level GPUs
- How to get started with visual generative AI on PCs with NVIDIA RTX
Price on eBay divides Tom’s Hardware readers
The weakest part of Oscar revenue is the vaunted cost of entry. In the comments on the article itself, readers contested: they report that the plate has been traded above $500 months ago.
Other users have linked listings for around US$140 for the bare board, with adapter and heatsink charged separately, which brings the total back closer to the original US$266. Dispersion is the hallmark of a second-hand market without standardized inventory.
The truth is that the rise in memory has pushed hardware prices up across the entire chain, and new video cards have already undergone adjustments of 10% to 50% in China as of July 25, a movement communicated by NVIDIA to its partners.
As long as local AI enthusiasts continue to soak up the remaining stock of Volta accelerators, the $100 per card window is likely to close on its own.
Sources): Tymscar, Tom’s Hardware and Wccftech
NVIDIA and AMD cards become more expensive in China as GDDR memory prices soar.
Source: www.adrenaline.com.br
Source link
