Should your business run its own AI model? We priced all three ways

By Greg Markowski / Jul 26, 2026 / Epic IT News

Every few weeks a client asks us some version of the same question: should we run our own AI model instead of paying per seat for Copilot or Claude? The pitch they have heard is seductive. Open weight models like Llama, Mistral, DeepSeek, and now Moonshot’s Kimi are free to download, your data never leaves your building, and the subscription line disappears from the P&L. So we priced it. All three ways a business can actually do this. Here is what the numbers say, and why the answer for most businesses is not the one the hype suggests.

What open weight models actually are

A quick definition, because the terms get muddled. An open weight model is an AI model whose parameters, the billions of numbers that make it work, are published for anyone to download and run. Meta’s Llama, Mistral, Alibaba’s Qwen, DeepSeek, and Moonshot AI’s Kimi are the big names. This is different from open source software in the traditional sense: you get the trained model, not always the data or the full recipe that produced it.

The commercial models you know, Claude, GPT, and Gemini, are closed. You access them through an API or an app, you pay per use or per seat, and the provider runs the infrastructure. The open weight pitch is that you take the model home and run it yourself. Where “home” is, and how good a model fits there, turns out to be the whole question.

Route one: your own hardware. The bill scales with ambition

Let us concede the cheap end first, because it is real. A Mac mini with enough unified memory will happily run a small quantised open model, and for one person drafting, summarising, and testing what these models can do offline, that is a genuinely useful setup for a few thousand dollars. If that is the ambition, buy the Mac and enjoy it.

That is not the question our clients are asking. They are asking whether a business can serve AI to forty or four hundred staff on its own hardware, and that is a different problem. A single small box serves roughly one user at a time at reading speed. Business serving means concurrency, response speed, uptime, and redundancy, and the hardware bill scales with how capable a model you want to serve. Mid-sized models on a serious GPU server start in the tens of thousands to buy, and dedicated cloud GPU capacity for stronger models runs thousands to tens of thousands of dollars a month around the clock.

And the top end just moved again. Moonshot AI’s Kimi K3, released this month as the strongest open weight model to date, needs roughly 1.4 terabytes of memory just to hold its compressed weights. That is multi-node GPU cluster territory, not a comms room. The pattern is the point: the open models weak enough to run on cheap hardware are noticeably behind the frontier models your staff compare everything against, and the open models that genuinely rival the frontier need infrastructure only a handful of operators own. Free weights, expensive electrons.

The governance bill is just as real as the hardware bill. Self-hosting makes you the AI vendor. The OAIC’s AI guidance is explicit that the Privacy Act and the Australian Privacy Principles apply to every use of AI involving personal information, including training, testing, and use. Self-host and the retention rules, access controls, and breach response a major provider carries contractually become entirely yours, and your cyber insurer will ask exactly how a server holding client data is patched and monitored. We covered what the Privacy Act changes mean for small business separately, and none of it gets lighter when the AI runs on your own tin. Most businesses that price route one only price the hardware.

There is a real niche where this maths inverts: organisations that cannot send data to an external API at all. Defence supply chain businesses working toward DISP obligations, some health environments, and operational technology sites with air-gap requirements. If that is you, self-hosting is a genuine conversation. For everyone else, it is an expensive way to get a worse answer.

Route two: open models in your own Azure tenancy

This is the version of “run your own AI model” that actually stacks up, and it is the one the vendors selling GPU servers do not mention. Microsoft’s Azure AI Foundry lets you deploy open weight models from its model catalogue inside your own Azure tenancy, in an Australian region, billed per token on the Azure subscription you already have.

Read that list again, because each item is one of the reasons people wanted to self-host in the first place. Your data stays inside your tenant boundary and inside Australia. There is no per-seat subscription. And there is no GPU server in your comms room, because the serverless deployment means you pay for what you use and nothing when you do not.

For a business that already runs on Microsoft 365 and Azure, which is most of the businesses we support, this is not exotic infrastructure. It is another workload in a tenancy you already govern, covered by the identity, conditional access, and audit controls you already have. That is why we treat it as an extension of managed IT rather than a science project.

Route three: managed frontier models. What we run ourselves

The third route is the one we chose for our own business: a frontier model, deployed and governed properly. We wrote up why a Microsoft partner chose Claude over Copilot, and eighteen months on the reasoning has held. Frontier models are simply better at the work, API pricing has fallen to the point where cost is rarely the deciding factor at SMB usage levels, and the vendor carries the infrastructure risk.

The honest comparison across all three routes looks like this.

  Own hardware Open model in your Azure tenancy Managed frontier model
Upfront cost A few thousand (one user) to six figures (business serving) Near zero Near zero
Ongoing cost High and fixed at business scale Per token, scales with use Per user or per token
Model quality Limited by your hardware Good, below frontier Best available
Data location Your premises Your tenancy, Australian region Vendor cloud, contractual controls
Who carries governance All yours Shared with Microsoft, run by your MSP Vendor terms plus your policy layer
Who fixes it at 2am You Microsoft platform, your MSP config The vendor
Right for Solo experimentation, air-gap and sovereignty mandates Data-sensitive work, variable usage Most businesses, most work

Most of our clients land on route three for general work, with route two as the answer for specific workloads where the data cannot leave their tenancy. Almost none should be on route one for business serving, and the few that should already know who they are.

The part nobody prices: governance and connection to your business

Here is what two years of running AI inside our own business taught us, and it is the part every “host your own model” pitch leaves out. The model is maybe a third of the value. An AI that cannot see your ticketing, job management, finance, or documentation systems is a very clever intern locked in a room with no files. And an AI that can see them without rules is a liability.

The work that matters is the layer between the model and the business: connections into the systems your team actually works in, per-user permissions so the model sees only what that person is allowed to see, audit logging of every action, and a written policy covering acceptable use and what the model may never touch. We built that layer for our own operations first, across our service desk, client records, finance, and documentation platforms, and it is the reason our answer to “which model” is honestly “it matters less than you think”. The model is swappable. The governance and integration layer is the asset.

This is also where the three routes stop being different products and become the same discipline. Whichever route you choose, the governance requirements do not change: an AI acceptable use policy your board has seen, access controls that mirror your existing permissions, logs you can hand an auditor or insurer, and a straight answer to the question every client and regulator is starting to ask, which is “what does your AI have access to and who approved it”. Businesses that spend the budget on hardware or deployment and skip this arrive at go-live with a model that is technically theirs and practically ungovernable. Price the governance and integration before you price the model.

What you should do now

Write down why you want to run your own AI model. If the answer is cost, the maths almost certainly says otherwise at your size. If the answer is data sovereignty, route two probably gives you what you actually need without the hardware. Only a written mandate that data cannot touch an external API justifies route one at business scale.

Audit what your staff are already using. In our experience the self-hosting question usually arrives at businesses whose staff are already pasting client data into free consumer AI tools. Our shadow AI audit playbook shows how to find them. Fix that governance gap first, because it is a live risk today, not a hypothetical one.

Get the options priced against your actual usage. Our AI readiness assessment maps your workloads, your data sensitivity, and your Microsoft licensing against all three routes, and gives you the comparison in dollars rather than vibes. If route two or three is right for you, our Managed AI service runs it end to end.

Frequently asked questions

What are open weight AI models?

Open weight models are AI models whose trained parameters are published for anyone to download and run on their own infrastructure. Llama, Mistral, Qwen, DeepSeek, and Kimi are the best known. They differ from closed models like Claude or GPT, which you can only access through the vendor’s API or apps.

Can I run an open weight AI model on a Mac mini?

Yes, smaller quantised models run well on Apple silicon with enough unified memory, and for one person it is a cheap way to experiment offline. It does not translate to business use: a single machine serves roughly one user at a time, and the stronger open models need far more memory than any desktop has. Kimi K3, the strongest open model released so far, needs around 1.4 terabytes just for its weights.

Is it cheaper to run your own AI model than pay for Copilot or Claude?

Almost never at small and mid-sized business scale. Serving capable models to a whole team needs GPU infrastructure that costs more per month than most businesses’ entire software bill, while frontier model API pricing has fallen sharply. Self-hosting only wins financially at very high, constant usage, or where a sovereignty mandate removes the API option entirely.

Can I run an open weight model in my own Azure tenancy?

Yes. Azure AI Foundry deploys open weight models from its catalogue inside your own tenancy, in an Australian region, billed per token with no GPU hardware to own. Your data stays inside your tenant boundary and under your existing identity and audit controls. For Microsoft 365 businesses this is the practical middle path between subscriptions and self-hosting.

What governance do you need before deploying any AI model?

Four things, regardless of which route you choose: an acceptable use policy your leadership has approved, access controls so the model only sees what each user is permitted to see, audit logging of every AI action, and a documented answer to what the model can access and who approved it. If you run your own AI model, you also inherit the Privacy Act obligations a vendor would otherwise carry.

Which businesses should actually self-host an AI model?

Organisations with a hard mandate that data cannot reach any external API: parts of the defence supply chain, some health environments, and air-gapped operational technology sites. For everyone else, running your own AI model on your own hardware means paying more for a weaker model you now have to operate like a production system.

Working out which AI route fits your business?

We run AI across our own operations, in Azure tenancies, and on frontier APIs, so the advice comes from operating all three. Book a free AI readiness assessment and get the comparison in dollars.

Book a Free Assessment

About the Author
Written by Greg Markowski, Founding Director of Epic IT, a CRN Fast50-recognised Microsoft Solutions Partner managing IT and cybersecurity for Perth businesses since 2003. Greg holds a Degree in Computer Science and a Diploma in Computer Systems Engineering from Edith Cowan University, and is ITIL certified.

Further Reading

Previous

AI cyber security threats: what ASD's warning means for you

Return to News
Back to News
Next

Lifeline data breach: five things every Australian charity should check this week