Deployment options
Three ways to run it. All of them stay yours.
The right one depends on how strict your data rules are and whether you would rather own hardware or rent it.
Deployment options
Three ways to run it. All of them stay yours.
The right one depends on how strict your data rules are and whether you would rather own hardware or rent it.
On-premise
A machine in your own rack or office.
Best for: Strictest data rules, offline sites, long horizons
- Runs on hardware you buy and own outright
- Works with no internet connection at all
- We manage it over a VPN or SSH channel you control
- Predictable cost after the hardware is paid for
Private cloud
Dedicated GPU hosting in an EU data centre.
Best for: Fast start, heavier models, no hardware purchase
- Single-tenant GPU, not a shared inference API
- EU region of your choosing, contract in your name
- Scale up or down as usage grows
- Live in days rather than after a procurement cycle
Air-gapped
No network path in or out. At all.
Best for: Classified, clinical, or regulator-facing environments
- Models and updates delivered on physical media
- No telemetry, no licence check, no phone home
- Full audit trail of every change we make
- On-site deployment and training
Sizing
What the machine looks like.
Three shapes cover almost every deployment. We pick between them during the assessment, on your real volumes rather than on a spec sheet.
Team
01One department, tens of users
Apple silicon workstation or a single mid-range GPU
- Document search and internal Q&A
- Drafting and summarising
- Quiet office, no data centre needed
Department
02Several teams, hundreds of users
Single professional GPU in a rack server
- Larger models at production latency
- Concurrent use across teams
- Room to grow without a rebuild
Enterprise
03Company-wide, thousands of users
Multi-GPU server, or several in a pool
- Frontier-class open models
- High concurrency with headroom
- Failover across nodes
Buy it, or rent the same shape from us in an EU data centre. The software is identical either way.
Questions
What your security team will ask.
01Does any of our data leave our network?
No. The model runs on hardware you control, and inference happens locally. There is no outbound call to a model provider. In an air-gapped deployment there is no outbound path at all, and we document the full data flow so your security team can verify it rather than take our word for it.
02Are local models good enough compared to the big cloud ones?
For the work most companies want, document search, summarising, drafting, classification, and internal Q&A, open models running on modest hardware are already strong enough to be useful every day. We test candidate models on your own documents during the pilot, so you judge quality on your work rather than on a benchmark.
03What hardware do we need?
It depends on model size and how many people use it at once. A small team assistant can run on a single workstation-class machine. Heavier workloads want a dedicated GPU server. We size it during the assessment and tell you honestly when renting a private GPU is cheaper than buying.
04How does this help with GDPR, NEN 7510, or DORA?
Keeping processing inside your own infrastructure removes the third-party processor, the international transfer, and the retention question in one move. We are not your compliance officer, but we give you the architecture documentation and audit logging your compliance officer needs to sign off.
05What does it cost?
There are two parts: a fixed-price engagement to design and deploy the system, and a running cost that is either hardware you own or a monthly private GPU. We quote both after the discovery call, before you commit to anything.
06What if we want to stop working with you?
You keep everything. The system runs on your infrastructure with open models and documented configuration, and your IT team gets the runbooks during handover. There is no licence to expire and nothing that stops working when we stop.
07Do you replace our IT team?
No. We do the part they have not spent the last two years on, model selection, retrieval, deployment, and tuning, and we hand it over in a shape they can operate. Most of our work is alongside an internal team, not instead of one.
08How quickly can we see something real?
A working pilot on your own documents usually takes two to three weeks after the architecture is agreed. The discovery call itself takes 45 minutes and costs nothing.
Next step
Tell us what your team is not allowed to do yet.
A 45-minute call. We look at your workflows, your data rules, and whether local AI is worth it for you. If it is not, we will tell you that instead of selling you a project.
Not ready for a call? Send the security brief to whoever has to approve it, or just reply to an email. All three work.