Run open LLMs on your own dedicated 🇪🇺 EU GPU server — fully private, no data sharing. Every model is benchmarked on real hardware, so we show you the GPU with the best price-performance and its measured speed.
Whether you already run a GPU server and want to serve a model on it with vLLM, or you're new and want to dive into private, EU-hosted LLM inference — this one-click Private LLM service is for you.
| Model | Kontekst | Parallel t/s | Enkel t/s | Timebaseret | Månedligt | Konfiguration |
|---|---|---|---|---|---|---|
| Llama 3.3 70B 4-bit Perfect on sparbox.m2 | 12,288 | 68 | 35 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3 30B 4-bit Perfect on sparbox.m2 | 84,992 | 877 | 168 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3.6 27B 4-bit Perfect on sparbox.m2 | 69,632 | 396 | 66 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3.6 35B 4-bit Perfect on sparbox.m2 | 62,464 | 747 | 143 | €0.79/h | €425.00/mo | sparbox.m2 |
| Gemma 4 31B 4-bit Perfect on sparbox.m2 | 17,408 | 390 | 62 | €0.79/h | €425.00/mo | sparbox.m2 |
| Gemma 4 E2B Perfect on sparbox.sm1 More context, higher speed on sparbox.m2 | 76,800 | 914 | 133 | €0.43/h | €230.00/mo | sparbox.sm1 |
| 117,760 +52% | 1,160 +27% | 184 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Gemma 4 E4B Perfect on sparbox.sm1 More context, higher speed on sparbox.m2 | 30,720 | 538 | 70 | €0.43/h | €230.00/mo | sparbox.sm1 |
| 95,232 +205% | 737 +37% | 104 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Granite 4.1 8B Perfect on novatesla.s1 More context, higher speed on sparbox.m2 | 41,984 | 247 | 39 | €0.33/h | €175.00/mo | novatesla.s1 |
| 76,800 +80% | 556 +125% | 76 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Llama 3.1 8B Perfect on sparbox.m2 | 95,232 | 621 | 86 | €0.79/h | €425.00/mo | sparbox.m2 |
| Phi 4 Perfect on sparbox.m2 | 14,336 | 367 | 50 | €0.79/h | €425.00/mo | sparbox.m2 |
| Phi 4 multimodal Perfect on sparbox.m2 | 117,760 | 941 | 144 | €0.79/h | €425.00/mo | sparbox.m2 |
| Ministral 3 14B Perfect on sparbox.sm1 More context, higher speed on sparbox.m2 | 20,480 | 429 | 56 | €0.43/h | €230.00/mo | sparbox.sm1 |
| 84,992 +300% | 623 +45% | 89 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Ministral 3 3B Perfect on sparbox.sm1 | 84,992 | 1,218 | 164 | €0.43/h | €230.00/mo | sparbox.sm1 |
| Mistral 7B Perfect on sparbox.sm1 | 28,672 | 399 | 53 | €0.43/h | €230.00/mo | sparbox.sm1 |
| Llama 3.1 8B 8-bit Perfect on ranger.s1 | 20,480 | 69 | 62 | €0.41/h | €220.00/mo | ranger.s1 |
| GPT OSS 20B Perfect on sparbox.m2 | 117,760 | 739 | 223 | €0.79/h | €425.00/mo | sparbox.m2 |
| Ornith 1.0 9B 4-bit Perfect on ranger.s1 | 52,224 | 78 | 59 | €0.41/h | €220.00/mo | ranger.s1 |
| Qwen3 Coder 30B 4-bit Perfect on sparbox.sm1 More context, higher speed on sparbox.m2 | 25,600 | 886 | 170 | €0.43/h | €230.00/mo | sparbox.sm1 |
| 84,992 +224% | 949 +7% | 167 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Qwen3 VL 32B 4-bit Perfect on sparbox.m2 | 37,888 | 429 | 63 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen2.5 32B 4-bit Perfect on sparbox.m2 | 28,672 | 438 | 64 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen2.5 7B Perfect on sparbox.sm1 | 28,672 | 400 | 51 | €0.43/h | €230.00/mo | sparbox.sm1 |
| Qwen3 14B Perfect on sparbox.m2 | 36,864 | 364 | 50 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3 32B 8-bit Perfect on sparbox.m2 | 18,432 | 297 | 42 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3 4B Perfect on sparbox.sm1 | 36,864 | 627 | 85 | €0.43/h | €230.00/mo | sparbox.sm1 |
| Qwen3 4B (2507) Perfect on sparbox.sm1 More context, higher speed on sparbox.m2 | 58,368 | 626 | 84 | €0.43/h | €230.00/mo | sparbox.sm1 |
| 117,760 +100% | 895 +43% | 121 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Qwen3 8B Perfect on sparbox.sm1 More context, higher speed on sparbox.m2 | 23,552 | 379 | 49 | €0.43/h | €230.00/mo | sparbox.sm1 |
| 36,864 +54% | 604 +59% | 83 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Qwen3 Coder 30B 8-bit Perfect on sparbox.m2 | 41,984 | 812 | 153 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3 VL 32B 8-bit Perfect on sparbox.m2 | 15,360 | 294 | 41 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3 VL 4B Perfect on novatesla.s1 More context, higher speed on sparbox.m2 | 69,632 | 436 | 68 | €0.33/h | €175.00/mo | novatesla.s1 |
| 117,760 +69% | 874 +101% | 119 | €0.79/h | €425.00/mo | sparbox.m2 | |
| Qwen3.5 9B Perfect on sparbox.m2 | 95,232 | 558 | 82 | €0.79/h | €425.00/mo | sparbox.m2 |
| Qwen3.6 27B 8-bit Perfect on sparbox.m2 | 28,672 | 164 | 47 | €0.79/h | €425.00/mo | sparbox.m2 |
| Gemma 4 31B 8-bit Perfect on sparbox.m2 | 8,192 | 286 | 43 | €0.79/h | €425.00/mo | sparbox.m2 |
Kontext = maksimal anvendelig kontekstvindue. Parallel/Enkelt = tokens/sec. Priser afspejler den bedste pris-præstation GPU pr. model. Priser er forudbetalt; du beholder fuld root-adgang.
You don't have to pick a model here — order a bare GPU server and run whatever you like on it, with full root access.
Bestil en GPU-serverSeconds after you hit Start, your private model is live behind an OpenAI-compatible endpoint — monitor throughput, context usage and requests from your own dashboard.
Trooper.AI giver dig en håndteret GPU-server forinstalleret med vLLM, klar til at servere ethvert åbent stort sprogmodel via en hurtig, OpenAI-kompatibel API. I stedet for at bruge timer på at vælge en GPU, installere CUDA-drivere, kompilere vLLM og justere startflag, vælger du blot et model ovenfor og klikker på Start. På få minutter får du en dedikeret EU-hospederet GPU-server med vLLM allerede kørende i den præcise konfiguration vi har benchmarket for det valgte model – herunder kontekstlængde, parallelisme og kommandolinjeargumenter inkluderet.
vLLM er den branchenormerede højtydende inference engine, men at få det klar til produktion kan være besværligt: at matche driver og CUDA-versioner, vælge --max-num-batched-tokens, --max-num-seqs samt den rette kontekstvindue for din GPUs VRAM, og sikre sig, at modellen faktisk indlæses. Vores administrerede GPU-server med forinstalleret vLLM fjerner dette arbejde. Alle konfigurationsmuligheder på denne side stammer direkte fra automatiserede benchmarktests på den egentlige hardware, så serveren du deplojer opfører sig præcis som testkørslen – ingen gætning.
Hver server kører på dedikerede bare-metal-GPU'er i ISO/IEC 27001-certificerede og GDPR-overholdende tyske datacentre. Du får fuld root-SSH-adgang, en vedvarende maskine samt et privat endpoint – dine prompts og data forlader aldrig din server. Da det er en hel GPU-server (ikke delt inferens), kan du også installere yderligere AI-programmer, finjustere eller køre billede- og lydmodeller sammen med dit LLM.
For hver model viser vi den GPU med bedst pris pr. ydeevne, inklusive ærlige timeløn og månedlige priser. Du kan gennemgå reelle eksempler på svar per model via Vis responskvalitet før implementering, så du ved både hastigheden og kvaliteten af svaret du betaler for. Betal timer til eksperimenter eller måneder til produktion – en administreret GPU-server forinstalleret med vLLM, der skaleres efter dine behov.
A dedicated private LLM GPU server gives you the whole machine: full root access, a private OpenAI-compatible endpoint, unlimited requests at a flat prepaid rate, and the freedom to fine-tune, swap models or run image and audio workloads alongside your LLM. Because the GPU is yours, throughput and context length are predictable and your prompts never leave your server — ideal for steady traffic and strict privacy requirements.
Pick a model from the matrix above and click Start: within minutes you get an EU-hosted GPU server with vLLM already running the exact benchmarked configuration — context length, parallelism and launch flags included. Pay hourly to experiment or monthly for production, and upgrade to a bigger GPU anytime for larger context or models without reinstalling.
Behind every private LLM GPU server is enterprise-grade, upcycled hardware maintained by our own team. Here, Markus and Jaimie are racking an NVIDIA A100 cluster in one of our ISO/IEC 27001-certified colocation data centers in Germany — the same class of GPU servers you deploy from this page. We upcycle high-performance components into optimized inference rigs, extending hardware lifecycles while reducing e-waste. We don't resell third-party capacity; we own and operate our own hardware in colocation data centers in Germany and the Netherlands, so we can guarantee performance, security, and data residency at every layer of the stack.
Your private vLLM server exposes an endpoint that is 100% compatible with the OpenAI Chat Completions API format (/v1/chat/completions). If your application already uses the OpenAI SDK — Python, Node.js, or any HTTP client — pointing it at your own server is a one-line change: update the base URL and API key. You get the same request and response schema, and full support for streaming, JSON mode, function calling, and multimodal inputs. No code rewrite, no new abstractions, no vendor lock-in — your integration stays portable and you stay in control.
Looking for an OpenAI API alternative hosted in Europe? A private LLM GPU server gives you equivalent Chat Completions API functionality with EU data residency, a flat prepaid price, and full ownership of the machine.
Because your server speaks the OpenAI Chat Completions API, you can plug it into virtually any AI tool, IDE or automation platform — just point the base URL at your own endpoint and add your key. A few popular examples:
… and many more — any app, agent or SDK that supports an OpenAI-compatible endpoint works out of the box.