I bought a datacenter GPU that doesn't fit in a normal motherboard, macgyvered the fan with jumper wires, and now I'm running a model that ties with Claude Sonnet 4.6 on benchmarks, all for £200.
Yes, that’s basically what the article is about. They run the LLM across both GPUs.
But that’s a feature of llama.cpp. SLI doesn’t really exist any more, and NVlink requires a specific setup, which the 4080 is not part of (the 3090 was the last consumer one, apparently). So you couldn’t pool the VRAM.
Yes, that’s basically what the article is about. They run the LLM across both GPUs.
But that’s a feature of llama.cpp. SLI doesn’t really exist any more, and NVlink requires a specific setup, which the 4080 is not part of (the 3090 was the last consumer one, apparently). So you couldn’t pool the VRAM.