How it works
From first call to production in about four weeks.
A fixed, boring sequence, an architecture with no black boxes in it, and hardware sized on your real volumes rather than a spec sheet.
How it works
From first call to production in about four weeks.
A fixed, boring sequence. You always know what happens next and what your security team needs to look at.
- 0145 min
Discovery call
We look at the work your team actually does and which parts are blocked today. If local AI is the wrong answer for you, we say so on this call.
- 02Week 1
Architecture and sizing
Model selection, hardware sizing, integration points, and a written data-flow diagram your security team can review and sign off.
- 03Weeks 2 to 3
Pilot deployment
We deploy on your hardware or a dedicated machine, wire it to a real document set, and put it in front of a small group of real users.
- 04Week 4
Rollout and handover
Access control, audit logging, monitoring, runbooks, and training. Your IT team can operate it. We stay on for support if you want us to.
Reference architecture
What we actually install.
No black box. Every layer is an open component your team can read, audit, and operate without us.
Access
Web UI and an OpenAI-compatible endpoint behind your own SSO
- OIDC / SAML
- Reverse proxy
- Per-team API keys
Orchestration
Retrieval, prompt assembly, tool calls, and per-request policy
- RAG pipeline
- Role-based scoping
- Full query audit log
Knowledge
Your documents, chunked and embedded, indexed on your own storage
- Vector store
- Connectors
- Incremental sync
Inference
Open-weight models served locally, pinned to a version you approve
- Model server
- GPU or unified memory
- Version pinning
Operations
Metrics, logs, alerting, and backups wired into what you already run
- Prometheus / Grafana
- Log shipping
- Restore drills
Deployed with Docker and systemd on Linux, or natively on Apple silicon. Configuration lives in your repo, not in our heads.
Sizing
What the machine looks like.
Three shapes cover almost every deployment. We pick between them during the assessment, on your real volumes rather than on a spec sheet.
Team
01One department, tens of users
Apple silicon workstation or a single mid-range GPU
- Document search and internal Q&A
- Drafting and summarising
- Quiet office, no data centre needed
Department
02Several teams, hundreds of users
Single professional GPU in a rack server
- Larger models at production latency
- Concurrent use across teams
- Room to grow without a rebuild
Enterprise
03Company-wide, thousands of users
Multi-GPU server, or several in a pool
- Frontier-class open models
- High concurrency with headroom
- Failover across nodes
Buy it, or rent the same shape from us in an EU data centre. The software is identical either way.
Next step
Tell us what your team is not allowed to do yet.
A 45-minute call. We look at your workflows, your data rules, and whether local AI is worth it for you. If it is not, we will tell you that instead of selling you a project.
Not ready for a call? Send the security brief to whoever has to approve it, or just reply to an email. All three work.