Skip to content
Novyant
All insights
  • AI Automation
  • Data Governance

Why the right AI model for a regulated business is probably a small one

The frontier model is the one in the headlines. The one that survives a compliance review is usually smaller, cheaper, fine-tuned on your own material, and running somewhere you can point at on a map.

By Roberto Sanson6 min read
A single small server on a concrete shelf inside an open bank vault

Every AI conversation starts in the same place: which model is best.

It is the wrong first question for most businesses, and it is especially wrong for the ones we work with, the insurers, manufacturers, institutions and clinical practices where the binding constraint is never raw capability.

The binding constraint is that somebody, at some point, is going to ask where the data went. And the answer "a third-party API in another country, and we are not certain what they retained" is a bad answer in a boardroom, a worse one in an audit, and an unrecoverable one in a dispute.

The interesting shift in 2026 is that you no longer have to choose between a model good enough to be useful and a deployment you can defend. For a large class of real tasks, you can now have both, and the model that gets you there is usually a small one.

What a small language model is, in plain terms

A small language model (SLM) is a language model with roughly a few million to around 7 billion parameters, small enough to run on ordinary hardware, a single server or sometimes a laptop, instead of a rented cluster.

The trade is straightforward. A frontier model knows an enormous amount about everything. A small model knows much less about everything, and can be taught to be very good at one thing: your document types, your terminology, your categories, your definition of a correct answer.

For general-purpose reasoning, the frontier model wins and it is not close. For "read this specific form and pull out these eleven fields, the way our senior adjuster would," the gap narrows sharply, and after fine-tuning on real examples, it frequently closes or inverts.

This is not a fringe position any more. TechCrunch's 2026 outlook quoted AT&T's Andy Markus predicting fine-tuned SLMs would become a staple of mature AI enterprises this year, on cost and performance grounds alone.

The three reasons it wins in a regulated environment

1. You can say where the data is

This is the one that decides procurement, and it is not primarily technical.

When a model runs on infrastructure you control, the answer to "where does our data go" is a location. You can name the machine, the jurisdiction and the retention policy, and you can put it in a contract without a vendor's sub-processor list in the middle of your answer.

That matters more in Mexico than it did two years ago. The March 2025 reform dissolved INAI and moved personal-data enforcement to a new body, Transparencia para el Pueblo, under the federal anticorruption ministry. The obligations did not disappear; the institution supervising them changed. In a period where the enforcement posture is genuinely unsettled, being able to demonstrate exactly where personal data is processed is worth more than a favourable reading of a clause.

For our clients this also shows up commercially rather than legally. A carrier's audit team asks the question. A university's committee asks it. The engagement proceeds or does not on the quality of the answer.

2. The economics invert at volume

A frontier API is cheap per call and expensive per year at scale. A small model you host is the reverse: real setup cost, then a marginal cost close to electricity.

The crossover point arrives faster than most teams expect, because the workloads worth automating are exactly the repetitive, high-volume ones: the same documents, thousands of times a month. That is the profile where per-call pricing compounds and a fixed-cost deployment stops looking conservative and starts looking obvious.

There is a latency argument too, and it is more operationally significant than the cost one. A local model answers in milliseconds without a round trip to another country, and it keeps answering when someone else's service has an incident.

3. You are not exposed to somebody else's roadmap

The risk that has become concrete in the last year is dependency. Models get deprecated, prices change, terms change, and a capability your operation now depends on can be repriced by a company you have no relationship with beyond a credit card.

This is the thesis behind a great deal of recent investment. Prime Intellect raised $130M in July 2026 specifically to help enterprises train and own their own models, at a $1B valuation, with customers including Ramp and Zapier. Ramp's co-founder reported their own agent beat frontier models on accuracy while running faster and far cheaper. Radical Ventures' David Katz framed the underlying anxiety plainly: how do you know your AI vendor will not eventually compete with you?

For a startup that is a strategic question. For a mid-sized insurer it is a simpler one: is the thing our claims operation depends on something we control?

Where a small model is the wrong answer

We would rather say this before a client discovers it.

Open-ended reasoning. If the task is "read this contract and tell us what is unusual about it," you want the biggest model available. Breadth is the whole job, and a specialised model will confidently miss the thing you needed.

Low volume. Standing up hosted inference to process forty documents a month is a hobby with a maintenance schedule. Use an API and move on.

No training data. Fine-tuning needs real examples with known-correct answers. If nobody has ever recorded what the right answer was, you have a data collection project first, and pretending otherwise wastes a quarter.

No one to operate it. A self-hosted model is infrastructure. It needs patching, monitoring and someone who is accountable when it stops. If that capacity does not exist and is not being hired, the honest recommendation is a managed service.

The realistic architecture for most organizations is not one or the other. It is a small model handling the high-volume, well-defined, sensitive work inside your boundary, and a frontier model called for the rare, hard, open-ended cases, with an explicit rule about what data is allowed to leave.

How to decide, in four questions

  1. How many times a month does this task run? Under a few hundred, use an API. In the thousands, model the fixed-cost option properly.
  2. Would you be comfortable reading this data aloud in a deposition? If not, it should not leave infrastructure you control.
  3. Do you have a hundred examples of the right answer? If yes, fine-tuning is available to you. If no, that is the first project.
  4. Who patches it at 2am? If there is no name, buy managed and revisit in a year.

The bottom line

The industry spent three years arguing that bigger is better, and for general intelligence that argument was correct.

Most business problems are not general intelligence problems. They are narrow, repetitive, high-volume, and wrapped in constraints about where information is allowed to sit. On that specific shape of problem, a smaller model running inside your own boundary is frequently more accurate, considerably cheaper, faster, and much easier to defend.

The question worth asking is not which model is best. It is which model you can still be running, unchanged and explainable, in three years.

We build both patterns depending on which the problem deserves. Our AI automation practice describes how we decide, and where AI automation pays covers the prior question of whether to automate the task at all.

Working on something like this?

We spend the first conversation understanding what you run on today. No pitch, and no obligation to build anything.

Book a call