
Hetzner, the bare-metal and cloud infrastructure provider, is experimenting with an OpenAI-compatible LLM inference API running on its own data centers. Currently free and experimental—with no SLA or production guarantee—it serves only the Qwen 35-billion-parameter model, but the real significance lies in Hetzner's potential to leverage its legendary hardware efficiency and low cost structure to build a competitive inference business. The next major test will be whether the company commits larger GPU clusters and higher-end hardware; if it does, it could become a serious challenger to existing inference providers.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Hetzner launched an experimental LLM inference service running an OpenAI-compatible API on its own hardware, currently serving only the Qwen/Qwen3.6-35B-A3B-FP8 model (a 35-billion-parameter Mixture-of-Experts model with 262K context window). The service has no billing, SLA, or production guarantee; Hetzner is testing whether users want it, how it scales, which features matter, and what load it can handle.
Why it matters
Hetzner, known for brutally efficient hardware operations and low-cost infrastructure, is exploring whether inference can become another low-margin commodity product—and potentially a way to keep spare GPU capacity busy and generate revenue. The company's existing strengths in buying, operating, and cost-efficiently managing hardware make it a plausible candidate to compete in open-weight inference, where margins are usually thin unless you have a cost or utilization advantage.
What to watch
The hardware question is critical. Hetzner's current public GPU lineup includes RTX 4000 SFF Ada (20 GB VRAM) and RTX PRO 6000 Blackwell Max-Q (96 GB VRAM)—fine for smaller models but not the dense multi-GPU systems needed for very large open models like GLM-5 (754 billion parameters). If Hetzner announces B200/B300-class hardware and a broader model catalogue, the experiment could become a serious competitive threat; if it stays at one or two smaller models, it remains a cool test but not a major player.
Hetzner is experimenting with LLM inference, offering an OpenAI-compatible API running on its own infrastructure. The move surprised the article's author, but it represents a thoughtful experimental approach: Hetzner explicitly states there is no billing, no SLA, no production guarantee, and only one model available. The company wants to learn whether users actually want the service, how the system scales, which features matter, and what kind of load it can handle.
The current offering is Hetzner Inference, which serves only the Qwen/Qwen3.6-35B-A3B-FP8 model—a 35-billion-parameter Mixture-of-Experts model with 3 billion active parameters. The model accepts text and images, has a 262K context window, and uses FP8-quantized weights, making it small enough to run without a massive GPU cluster while still useful for real workloads. Users create an API token in the Experiments dashboard, point an OpenAI client at Hetzner's base URL, and use it like any other inference API. Hetzner also published a tutorial for connecting OpenCode to the API without requiring custom code.
When tested on July 23, 2026, the API showed strong performance: 153 ms median time to first token across seven short requests on an already-open connection, and 224 output tokens per second across five longer generations capped at 512 tokens. The model itself behaved roughly as expected—it followed most formatting and retrieval instructions, handled images correctly, and failed on two very simple arithmetic questions. The author notes that an undocumented enable_thinking option helped prevent the model from spending too much of its completion budget on reasoning before returning a visible answer, though it is not yet documented by Hetzner.
The real strategic interest, however, lies not in the Qwen model but in what Hetzner might be building. Open-weight inference is a commodity market where anyone can download the same weights and expose an OpenAI-compatible API, making switching providers easy with tools like OpenRouter or LiteLLM. Building large margins requires either very cheap hardware procurement and operation, exceptional skill at keeping hardware busy, or spare GPU capacity that would otherwise sit idle. Hetzner excels at all three: the company's entire business model revolves around buying hardware, putting it into its own data centers, and operating it with brutal efficiency. An inference API could let Hetzner share GPU capacity across many users and keep hardware busy, turning spare capacity into revenue.
The critical question is hardware. Hetzner's current public GPU lineup includes RTX 4000 SFF Ada with 20 GB VRAM and RTX PRO 6000 Blackwell Max-Q with 96 GB VRAM—capable workstation GPUs suitable for smaller models. The FP8 files for the current Qwen model are around 38 GB, with actual VRAM use landing higher depending on context length and cache size. However, these are not the dense multi-GPU systems needed for very large open models. GLM-5, a 754-billion-parameter model, requires eight GPUs and hundreds of gigabytes of VRAM even with aggressive quantization—realistically B200/B300-class hardware with very fast links between GPUs. Hetzner does not currently offer such systems in its public bare-metal lineup. The article's author notes that internal hardware behind the experimental API may differ from what Hetzner sells publicly, but this gap is the critical signal to watch. If Hetzner keeps serving only one or two smaller models, it remains an interesting experiment. If the experiment is the first step toward larger GPU clusters, a proper model catalogue, and B200/B300-class hardware, Hetzner—with its data centers, network, hardware expertise, European positioning, and reputation for aggressive pricing—could become a serious inference competitor.
Hetzner's inference experiment is notable not because of the Qwen model itself—which the author describes as functional but merely a reasonable starting point—but because it reveals a plausible strategy for the company to enter a commodity market. Open-weight inference margins are thin because anyone can download the same model weights and run similar serving software; the main advantages come from either cheap hardware procurement and operation, keeping hardware busy with high utilization, or repurposing idle capacity. Hetzner excels at all three: it is famous for buying and operating GPU hardware with ruthless efficiency, and an inference API would allow it to share GPU capacity across many users rather than dedicating it to a single customer, potentially keeping expensive hardware in use full-time.
The critical constraint is hardware. Hetzner's public lineup includes workstation-class GPUs (RTX 4000 SFF Ada with 20 GB VRAM and RTX PRO 6000 Blackwell Max-Q with 96 GB VRAM) suitable for smaller models but not the dense multi-GPU clusters required for very large open models. For context, the GLM-5 model (754 billion parameters) requires eight GPUs and hundreds of gigabytes of VRAM even with aggressive quantization—technology at the B200/B300 level with fast multi-GPU interconnects. The author notes that Hetzner's internal infrastructure behind the experimental API might differ from its public catalogue, but the absence of such hardware in the public lineup is worth monitoring.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime