No, the requirements to run a locally hosted flagship llm is on the order of a Terabyte of vram, like Deepseek v4. The best models from openai and anthropic are likely similar. These are usually loaded into vram on one machine that can serve many users
But if you tried to distribute the compute, if it’s even possible, would take like a hundred high end gaming computers and process at a glacial pace. It just doesn’t make sense.
It makes more sense for an organization to purchase their own hardware and distribute resources somehow, like for a University, but consumer hardware won’t really cut it.
You could use smaller models, but their performance scales with size so they are much dumber than the flagship models.
No, the requirements to run a locally hosted flagship llm is on the order of a Terabyte of vram, like Deepseek v4. The best models from openai and anthropic are likely similar. These are usually loaded into vram on one machine that can serve many users
But if you tried to distribute the compute, if it’s even possible, would take like a hundred high end gaming computers and process at a glacial pace. It just doesn’t make sense.
It makes more sense for an organization to purchase their own hardware and distribute resources somehow, like for a University, but consumer hardware won’t really cut it.
You could use smaller models, but their performance scales with size so they are much dumber than the flagship models.