Sovereign inference hosting, sized to your models.
Dedicated model serving on Swiss soil. Every deployment is individual - model sizes, throughput, SLAs and quotas are shaped around your workload, so we build and price it with you rather than off a menu.
Built around your needs, not a hosting catalogue.
Inference serving is more than renting a GPU. We give you what a standard hoster does not:
Serve models the size you actually need.
A single dedicated allocation scales up to four GPUs of pooled accelerator memory. What it holds depends on the precision you run - larger models fit as you quantise. The figures below are the largest model a full allocation serves at each precision; smaller dedicated allocations are available, and we size every deployment to your model rather than the other way round.
| Precision | Largest model served |
|---|---|
| Full precision (FP16 / BF16) | up to ~230B parameters |
| FP8 | up to ~450B parameters |
| 4-bit (INT4) | up to ~700B parameters |
Indicative. Real capacity depends on context length, batch size and throughput targets - we confirm the exact fit with you before deployment.
A serving platform, run for you.
About isolated virtual servers: they run on a multi-tenant chassis. The compute host, BMC, power and PCIe fabric are shared, and isolation from co-tenants is enforced at the VM, IOMMU and network layer - not by physical dedication. Other tenants run on the same chassis in separate virtual servers, disclosed and accepted before delivery. As the operator of both the virtualisation layer and the serving stack, Silex Radix can technically reach the virtual server holding your models, including its running memory. Where your trust model must exclude the operator, choose the dedicated single-tenant option.
Let us size your inference deployment.
Bring your models and your throughput targets. We bring the sovereign Swiss capacity and run the stack.
Talk to us