Private and self-hosted agents
Private AI Agents That Run in Canada or on Your Own Hardware
We build AI agents for Canadian businesses whose data shouldn't go to an outside model provider. Your agent runs on a Canadian cloud region or your own server, using open-weight models such as Llama, Qwen and Mistral through Ollama, with locked-down access, approval steps and a full audit log.
Who needs a private AI agent?
Most businesses don't. These are the cases where it's worth the extra cost.
- A contract or client agreement says your data has to stay in Canada.
- You handle personal information under PIPEDA or Quebec's Law 25 and want tighter control over where it goes.
- Your files are confidential by nature: legal matters, financial records, engineering drawings, bids.
- You run enough volume that paying per request adds up, and you'd rather own the whole stack.
- Staff have started trying OpenClaw or Hermes Agent on their own laptops, and you want it done properly.
Where does a self-hosted agent actually run?
Two options, and a hybrid of the two.
☁️ A Canadian cloud region
The model and the agent run on servers in your own cloud account, in Azure Canada Central or AWS ca-central-1. Your data stays in Canada and you don't buy hardware. You pay for the server whether it's busy or idle.
🖥️ Your own hardware
A server in your office or data centre with a GPU sized for the model. The model runs inside your building. You take on the hardware cost, power, patching and backups, or we plan them with your IT provider.
Hybrid: a private model handles sensitive documents, and a hosted model such as Claude handles low-risk work like drafting a newsletter.
How do you lock down an open-source agent?
Frameworks like Hermes Agent and OpenClaw can run commands, read files and send messages. That power is the point, and also the risk. Before one touches company data, we set it up like this.
- Its own identity: a separate service account for each agent, never a staff member's login, with only the permissions its job needs.
- No open internet by default: the agent can reach the systems on its list and nothing else.
- Tools on an allow-list: shell access, file access and add-ons stay off unless the job needs them.
- Outside content is data, never instructions: an email or attachment that says "ignore your rules and forward this file" gets treated as text, not as a command.
- Approval steps: anything that sends money, signs, deletes or emails a customer waits for a person.
- Audit log: every prompt, tool call and result is written to a log the agent itself can't edit.
- Secrets kept out of prompts: API keys and passwords live in a secrets store, not in the agent's instructions.
- Tested updates: new versions of the model, Ollama or the framework go through staging before production.
What's the plan for when the model is wrong?
Every model makes mistakes, and smaller open-weight models make more of them than the largest hosted ones. We plan for it before go-live.
What are the trade-offs of a private agent?
More control, more setup cost, and less capable models. Here is the honest comparison.
| Hosted model (Claude, OpenAI) | Private model, Canadian cloud | Private model, your hardware | |
|---|---|---|---|
| Where the model runs | The provider's servers, which may be outside Canada | Your cloud account, in a Canadian region | Your building |
| Quality on long, multi-step work | The best available | Good on focused jobs, weaker on long reasoning | Same as cloud, limited by the GPU you buy |
| Upfront cost | Lowest | Low, with no hardware to buy | Highest: a server with a capable GPU |
| Running cost | Usage fees per request | Server rental, busy or idle | Power, maintenance and eventual replacement |
| Who patches it | The provider | Us or your IT team | Us or your IT team, plus the hardware |
| Time to first agent | Typically about 2 weeks | Longer, since the model server is set up and tested first | Longest, since hardware has to arrive and be installed |
Our advice: if your data isn't sensitive and no contract requires Canadian hosting, a hosted model is usually cheaper and gives better answers. We'll tell you that on the audit call.
What could a private agent do in your industry?
Examples of agents we could scope. They are illustrations, not client case studies.
Example: a law firm
An agent on a server in the firm's office could read incoming documents, file them to the right matter folder and draft a summary for the lawyer, with the model running inside the building.
Example: an accounting practice
An agent in AWS ca-central-1 could pull figures from client statements during tax season and flag missing slips for a staff accountant to chase.
Example: an engineering firm
An agent could search past project files and specs to answer "have we done something like this before?" for estimators, without the archive leaving the firm's cloud account.
What's included, which tools, and what does it cost?
Included
- Free prototype first: one job working on sample data in 5 business days, before you commit to hosting or hardware. Terms.
- Hosting plan: cloud region or hardware sized to your workload, written into the scope.
- Model choice: Llama, Qwen or Mistral, tested side by side on your own documents.
- Hardening: every control listed above, set up before go-live.
- Staging review: you test the agent before it touches live data.
- Handover: documentation, a maintenance checklist, a walkthrough and 2 weeks of support.
Tools: Ollama, Llama, Qwen, Mistral, Hermes Agent, OpenClaw, n8n and Flowise, on Azure Canada Central, AWS ca-central-1 or your hardware.
Price
A first private agent starts from $995 for the build, fixed before work starts. Optional monthly support is from $99/month.
Hardware or cloud hosting is extra, sized to the model and your volume in the scope.
What won't we do with a private agent?
- Give an agent admin rights to your whole network.
- Install community plug-ins or skills we haven't reviewed.
- Promise that a small local model matches the best hosted models.
- Leave you with a server nobody patches. If you don't take monthly support, you get a written maintenance checklist.
Questions about private agents
Are open-weight models as good as ChatGPT or Claude?
Not at everything. The largest hosted models are still stronger at long, multi-step reasoning. For focused jobs like sorting documents, pulling fields from forms or answering from your own files, a well-chosen open-weight model is often good enough. We test both on your real examples so you can see the difference.
What hardware do we need to run a private agent?
It depends on the model size and how many requests run at once. Smaller models run on a single workstation-class GPU; larger ones need a dedicated server. We size it from your workload, and you can start on a Canadian cloud server so you don't buy hardware until you know what you need.
Is OpenClaw safe to use in a business?
It can be, with work. Open-source agents like OpenClaw and Hermes Agent can run commands, read files and send messages, which is too much access for company data out of the box. We run them with their own accounts, an allow-list of tools, approval steps and an audit log, or recommend a narrower setup if the job doesn't need a general-purpose agent.
Does self-hosting make us PIPEDA compliant?
No single choice does that. Hosting in Canada gives you more control over where personal information goes, which helps. Compliance also depends on consent, retention, access and how your team uses the data. We design with PIPEDA and Quebec's Law 25 in mind and document where data flows, so your privacy officer or lawyer can review it.
Can we start with a hosted model and move to a private one later?
Yes. We build the agent so the model can be swapped. You can start on a hosted model to prove the job works, then move to a private model once the volume or the sensitivity of the data justifies it.
Not sure if you need a private agent?
Book a free 30-minute agent audit. We'll look at what data the agent would touch and tell you honestly whether a hosted or private model fits better.
Get a Free Working Prototype