Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU
This is the vLLM entry in my local AI series. After testing Ollama, llama.cpp, and FreeToken, I wanted to run vLLM as a Kubernetes Deployment on my Windows/WSL2 setup. The plan was simple: deploy vLLM, request nvidia.com/gpu: 1, expose an OpenAI-compatible API endpoint through a Service, and tie it into the Kubernetes workflows I write about regularly. Getting a GPU into Kubernetes on WSL2 turned into an investigation. Not because vLLM is hard, but because container runtimes and nested clusters handle GPU passthrough in non-obvious ways. Here is what happened, how the device plugin and node prerequisites work, and how to get a working GPU cluster running with Minikube. ...