Keera Engine

Open models, run efficiently

Keera Engine is the inference engine behind the gateway. It runs any open model on open technology - Kubernetes, vLLM and llm-d - on Swiss GPUs or in your own cluster. More answers per GPU, and no layer is a black box.

Try Keera See Keera Gateway ➔

Keera coding on her laptop

What Keera Engine is

An inference engine built from open parts: Kubernetes schedules the GPUs, vLLM and llm-d serve the requests, open models run on top. Above it all sits an ordinary API - whatever talks to OpenAI today talks to Keera tomorrow.

The API

OpenAI-compatible. One base URL, one key, and any client that already speaks to OpenAI speaks to Keera - no new SDK, no plugin.

The models

Model-agnostic: any open model vLLM can serve, plus three Keera fine-tunes and your own. You can read what is inside them and take all of it with you whenever you like.

The inference server

vLLM batches many requests onto one GPU. llm-d sends each one to where its context already sits in cache. The same hardware computes more tokens.

The platform

Kubernetes places the models on the GPUs and scales them with the load. In Swiss data centres or in your own racks, with no internet at all if that is what you need.

Top to bottom: what your client sees, and what runs underneath it.

The models

The engine runs any open model vLLM can serve - Qwen, Llama, Mistral, Apertus or your own fine-tune. Three fine-tunes of ours run from day one. Pick one per workspace - or let the gateway route: small tasks to the fast model, hard reasoning to the strong one.

Model Context Best for
Keera Deep 256k Agentic work across many files
Keera Swiss 128k Code and documentation in German, French and Italian
Keera Fast 64k Completion, tests, commit hygiene
Your own fine-tune - Internal frameworks and legacy code
Any other open model - Whatever your use case needs

Keera Deep and Keera Fast are derived from Qwen-Coder, Keera Swiss from Apertus. All three are Apache-2.0: you can download the weights and take them with you.

Hosted providers - today the OpenAI models - sit behind the same policies, but run on the provider's infrastructure and not on Swiss GPUs. They are blocked by default; a tenant opens them per team and data class, or not at all. Checked September 2026.

Open technology, run by us

Kubernetes, vLLM and llm-d are open source and widely used. We reinvent none of it. We tune the parts to each other and run them - in Swiss data centres, or together with your team in your own cluster.

Every token the engine computes passes through Keera Gateway first: the same policies and the same audit log as everywhere else.

Keera and an engineer at the same screen

Common questions

What platform and engineering teams ask us most before they try Keera Engine.

Does my data leave the building?

No. The models run on Swiss GPUs or in your own racks; prompts and code go there for inference and come back. They are not stored, the model's answer is not stored, and nothing is trained on either. What is kept is the request itself - who, when, which model - in the log in your own infrastructure.

Are open models good enough?

For everyday work, yes: refactoring across many files, tests, migrations, reviews. On the hardest reasoning the large proprietary models are still ahead - which is why they sit behind the same gateway and can be opened per team and data class. We do not claim open weights win everywhere. We claim you keep the choice.

How does this fit with Keera Gateway?

Keera Engine sits behind the gateway and runs the open models the gateway routes to. The gateway decides who may use which model and writes every request to the audit log; the engine computes the answer. You can start with the gateway and the providers you already have, and add the engine once the models should move in-house.

Try Keera

We put one team on the engine, measure throughput and cost with you and hand you the numbers. After 30 days you decide.

Interested in
We use your details only to answer this enquiry.