Inference hosting · Switzerland

Sovereign inference hosting, sized to your models.

Dedicated model serving on Swiss soil. Every deployment is individual - model sizes, throughput, SLAs and quotas are shaped around your workload, so we build and price it with you rather than off a menu.

Built around your needs, not a hosting catalogue.

Inference serving is more than renting a GPU. We give you what a standard hoster does not:

Tuned to how your team works
Caching and load balancing shaped to your workload, so answers start fast and stay fast as more of your people use the service.
Long work picks up where it left off
A paused analysis, a long document or a multi-step agent resumes without re-reading everything from the start.
Voice agents
Connect voice agents - speech in, speech out - to the models you serve.
Bring your own models
Your fine-tunes or the open-weight model you choose, served and tuned for your use rather than picked from a fixed menu.
Capacity that follows real usage
We look at how you actually use the service before adding hardware, so more users do not automatically mean a bigger bill.

Serve models the size you actually need.

A single dedicated allocation scales up to four GPUs of pooled accelerator memory. What it holds depends on the precision you run - larger models fit as you quantise. The figures below are the largest model a full allocation serves at each precision; smaller dedicated allocations are available, and we size every deployment to your model rather than the other way round.

Largest model served per precision at a full dedicated allocation - indicative.
Precision Largest model served
Full precision (FP16 / BF16) up to ~230B parameters
FP8 up to ~450B parameters
4-bit (INT4) up to ~700B parameters

Indicative. Real capacity depends on context length, batch size and throughput targets - we confirm the exact fit with you before deployment.

A serving platform, run for you.

Self-service model portal
Switch the models you serve yourself - manually from the portal, or automatically on a schedule or by demand.
Multi-model serving
Run several models side by side behind one endpoint, each isolated in its own process.
Batch processing
High-throughput offline jobs - large document sets, embeddings, evaluations - on the same dedicated capacity.
OpenAI-compatible API
A drop-in /v1 endpoint. Point your existing SDKs and agents at it and keep your code.
Integrations on request
Connect the deployment to what you already run - your own object storage and data buckets, retrieval over your own corpus (RAG), and your internal tooling. Scoped and quoted per deployment rather than sold as a fixed feature list.
Dedicated and single-tenant
Your own hardware allocation on Swiss soil. No co-tenancy, no US-jurisdiction exposure, and requests never leave the boundary.
Isolated virtual servers
A lower-cost option, with the same portal and the same /v1 endpoint: your models are served from a virtual server of their own, with an enterprise GPU passed through to it alone and hardware-enforced isolation from the other tenants on the chassis. You consume it as a service, we operate it. Read the note below for what this tier does and does not guarantee.
Sandboxed and self-contained
Each model runs in its own sandboxed process, with no third party and no subprocessor in the request path. Your content stays yours - we process it only as far as running your service requires, never to train anything, never across tenants. Silex Radix GmbH is your processor under a Swiss data-processing agreement (nDSG): you remain the controller, and we notify you of any security incident.
Your models, your weights
Run open-weight models or your own adapted checkpoints. We serve what you bring, and nothing is shared across tenants.
Autoscaling and quotas
Scale models up and down within your allocation, with rate limits and quotas set to your plan.
SLAs and observability
Agreed uptime and throughput targets, with usage metering and monitoring you can see.

About isolated virtual servers: they run on a multi-tenant chassis. The compute host, BMC, power and PCIe fabric are shared, and isolation from co-tenants is enforced at the VM, IOMMU and network layer - not by physical dedication. Other tenants run on the same chassis in separate virtual servers, disclosed and accepted before delivery. As the operator of both the virtualisation layer and the serving stack, Silex Radix can technically reach the virtual server holding your models, including its running memory. Where your trust model must exclude the operator, choose the dedicated single-tenant option.

Priced to your deployment, not a price list.
What you pay tracks what you run - model sizes, throughput, how many models, dedicated versus burst capacity, and the SLAs and quotas you need. Rather than force that into a standard tier, we work it out with you: tell us your workload and we shape a deployment and a price around it.

Let us size your inference deployment.

Bring your models and your throughput targets. We bring the sovereign Swiss capacity and run the stack.

Talk to us