$266 Tesla V100 PC runs AAA games and 27 billion parameter AI model

Emphasis Hardware

A Tesla V100 purchased for around US$200 became a video memory expansion in a gaming PC that was already running AAA games with an RTX 4080. The set included 32 GB of VRAM and started running a language model with 27 billion parameters at 32 tokens per second, without depending on the cloud or internet connection.

The editing is by Oscar Molnar, who documented the process in your blog however, he emphasizes that GeForce is still the one who designs the frames. The server board came in only to load the model into memory.

What cost $266

The total disbursement was £200, something close to US$266. The approximate division looks like this, with conversion at the commercial rate of R$5.08 this Monday (27):

Item Dollar value Approximate conversion
Tesla V100 SXM2 16GB Around US$200 R$ 1,016
SXM2 to PCIe Adapter About US$66 R$ 335
Total $266 R$ 1,351

The board does not have a PCIe slot, video output or power connector

The Tesla V100 used is the SXM2 variant, a format that NVIDIA created for modules slotted directly into server baseboards. It has no PCIe contact edge, no PCIe power connector and no video output.

Reproduction/Oscar Molnar

Installing this card in a home machine requires an SXM2 to PCIe adapter, a part sold on the black market for around US$66. It is the item that translates the server socket into the x16 slot on a common motherboard.

There is a 32 GB version of the same card, which would double the memory in one go. The price, however, also doubles, and it circulates in the US$400 to US$500 range on the second-hand market.

82 dB and the fan mod

The original V100 SXM2 cooler was designed for racks, where no one can hear anything. Measured in a home environment, it recorded 82 dB. Molnar described the result as something “between a garbage disposal and a lawnmower.”

Reproduction/Oscar Molnar

The solution was simple and did not require an expensive part, the fan wires were redirected to the motherboard’s PWM header, which allowed the operating system to control the rotation. It also works to buy a 2.54 mm male to PH2.0 female jumper cable and bridge it.

With rotation control available, the fan started to rotate at 10% of maximum speed. Even at this level, the card remains below 50 °C under full load, which provides a comfortable margin for prolonged operation.

How Volta and Ada live on the same machine

The other obstacle was software: the PC combines two architectures separated by five years: the RTX 4080 is Ada Lovelace, launched in 2022, and the Tesla V100 is Volta, from 2017.

The solution was to run NixOS with a legacy NVIDIA driver, a version whose compatibility window still covers both architectures at the same time. From there, the system sees a combined 32 GB of VRAM and distributes the model layers between the two cards for local inference.

The V100 continues to deliver respectable numbers almost a decade after launch. There are 5,120 CUDA cores and a 4,096-bit bus that supports around 900 GB/s of bandwidth, a level that most consumer GeForce devices do not reach.

Specification Tesla V100 SXM2
Architecture Volta (GV100)
CUDA Cores 5,120
Memory 16GB HBM2
Memory bus 4,096 bits
Bandwidth Around 900 GB/s
Tensor Cores First generation
Ray Tracing Cores None
Video output None
Physical interface SXM2, requires adapter

The absence of cores dedicated to Ray Tracing and any video output explains why the card does not replace a GeForce.

In language model inference, however, the math changes: the task is limited by memory bandwidth, and that is precisely where HBM2 has an advantage over GDDR6 on newer mid-range cards.

Reproduction/Oscar Molnar

Qwen3.6-27B at 32 tokens per second

The test ran Qwen3.6-27B-MTP quantized in Q5_K_M, 19 GB file, with a context window of 128 thousand tokens. There was memory left for the entire context without having to use the system’s RAM.

Metric Result
Model Qwen3.6-27B-MTP
Quantization Q5_K_M
File size 19GB
Context window 128 thousand tokens
Text generation 32 tokens per second
Prompt processing 133 to 160 tokens per second
Total system VRAM 32 GB (16 GB from RTX 4080 plus 16 GB from V100)

“Fast enough for interactive use”, summarized Oscar Molnar in the report published on his blog, adding that the rate obtained exceeds that of most alternatives accessed via API in the cloud.

For comparison, comfortable human reading is around 10 to 15 tokens per second. The 32 tokens per second puts the response above the reading speed, which eliminates the feeling of waiting while typing output.

Also read:

  • 8-year-old NVIDIA V100 GPU becomes AI phenomenon after dropping to $100
  • Micron launches 3GB GDDR7 memory to increase VRAM in entry-level GPUs
  • How to get started with visual generative AI on PCs with NVIDIA RTX

Price on eBay divides Tom’s Hardware readers

The weakest part of Oscar revenue is the vaunted cost of entry. In the comments on the article itself, readers contested: they report that the plate has been traded above $500 months ago.

Other users have linked listings for around US$140 for the bare board, with adapter and heatsink charged separately, which brings the total back closer to the original US$266. Dispersion is the hallmark of a second-hand market without standardized inventory.

The truth is that the rise in memory has pushed hardware prices up across the entire chain, and new video cards have already undergone adjustments of 10% to 50% in China as of July 25, a movement communicated by NVIDIA to its partners.

As long as local AI enthusiasts continue to soak up the remaining stock of Volta accelerators, the $100 per card window is likely to close on its own.

Sources): Tymscar, Tom’s Hardware and Wccftech

Related Content
NVIDIA and AMD cards become more expensive in China as GDDR memory prices soar

It’s not easy for anyone!

NVIDIA and AMD cards become more expensive in China as GDDR memory prices soar.

Source: www.adrenaline.com.br
Source link

Leave a Reply

Your email address will not be published. Required fields are marked *

thirteen − four =