When self hosting makes sense
- You handle data that should not leave your network: health, legal, financial, personal or classified.
- Your contracts or regulators limit where data can be processed.
- Your usage is steady enough that owning hardware beats paying per user or per token.
- You want to keep working whatever happens to a provider's prices, terms or models.
It makes less sense when you need the very strongest model for open ended work, your usage is occasional, or nobody can own the system day to day.
The stack
| Layer | What it does | Common open source options |
|---|---|---|
| Hardware | GPUs with enough memory for the model and its users | One workstation GPU to a multi GPU server |
| Inference server | Loads the model and serves it through an API | vLLM for many users; llama.cpp or Ollama for small setups |
| Interface | Chat and document tools for staff | Open WebUI and similar chat front ends |
| Retrieval | Search over your documents with citations | A vector database such as pgvector or Qdrant |
| Identity | Who can use what | Your existing single sign on, with MFA |
| Observability | Usage, errors, performance and audit logs | Your existing logging and monitoring stack |
Sizing the hardware
GPU memory is the constraint that matters most. A model's weights need roughly its parameter count times the bytes per parameter: about 2 bytes at 16 bit precision, and about half a byte at 4 bit quantisation. On top of that, every active conversation needs memory for its context. As a rough guide:
- A 30 billion parameter model quantised to 4 bits needs about 15 to 20 GB for its weights, which fits a single 24 to 32 GB workstation GPU with room for some context.
- Long documents and many simultaneous users multiply the context memory, so serving a team needs more headroom than a single user.
- Mixture of experts models activate only part of their weights per token, so they run faster than their total size suggests, but still need memory for all of it.
Measure with your own model, context length and number of users before you buy. Quantisation saves memory at some cost in quality, so test the quantised model on your real tasks.
Choosing a model
- Licence. Apache 2.0 and MIT licences allow commercial use with few conditions. Some open weight models use custom licences with usage restrictions or attribution rules. Read the licence before you build on a model.
- Fit. Pick the smallest model that does your task well. Retrieval heavy work often needs less model than you'd expect, because the answer comes from your documents.
- Evaluate on your work. Public leaderboards are a starting point. Build a small test set from real tasks and compare models on it.
- Local defaults. Check how the model handles Australian spelling, law and context for your use case.
Security controls that matter
A self hosted AI system is an information system, and the Essential Eight and ISM controls still apply. These are the controls we found mattered most when building on premises AI platforms.
Network and application
- Keep the inference server off the internet; expose only the application, behind TLS 1.2 or higher with HSTS.
- Put a web application firewall in front, and use a strict content security policy.
- Make sure the model itself makes no outbound calls unless you intend it to.
Identity and access
- Use single sign on with MFA, and role based access to models, tools and document collections.
- Carry document permissions into retrieval, so each person can only retrieve what they could open.
- Keep retrieval indexes and conversation history isolated per user or per team.
Model specific risks
- Keep system instructions separate from user input and retrieved content, to limit prompt injection.
- Sanitise model output before rendering it, and never execute it without checks.
- Give models and agents the fewest tools and permissions they need.
- Rate limit users and cap the length of requests and responses.
The OWASP Top 10 for LLM Applications covers these risks in depth.
Data and operations
- Encrypt stored conversations and documents, and set a retention period for them.
- Keep an audit trail of who used what, when.
- Patch the operating system, drivers, inference server and dependencies on a schedule.
- Back up configuration, indexes and conversation data, and test restoring them.
- Have an incident response plan that covers AI specific incidents.
Running it well
Pin model versions and test before upgrading, because a new model can change behaviour your staff rely on. Watch usage and response times to know when to add capacity. And keep a small evaluation set you rerun after every change, so improvements are measured rather than assumed.
Checklist
- A use case with a named owner and a measure of success
- Hardware sized for your model, context length and users
- A model whose licence fits your use
- SSO, MFA and role based access
- No public exposure of the inference server
- Prompt injection, output handling and tool limits addressed
- Audit logs, retention, backups and patching in place
- An evaluation set to test every change
This guide is general information, not legal or professional advice. Check the official sources, and get advice for your situation.
More guides:
- AI governance: A practical framework for adopting AI responsibly.
- Australian AI compliance: The laws, standards and guidance to know.
- Data sovereignty: Where your data goes when you use AI.