Skip to content
Novyant
All insights
  • Data & Analytics
  • AI Automation

Your AI project is a data project, and most of the work happens first

Almost every AI request we receive resolves, within an hour of looking, into four ordinary data problems. The good news is that you do not need clean data. You need the specific data this one decision depends on.

By Ciro Calderon5 min read
An archive drawer of index cards with one card lifted clear of the row

We get asked for AI. We almost always find a data project.

That sounds like a consultant's deflection, so here is the specific claim: in most engagements, the thing preventing a useful AI system is not model capability, budget or engineering capacity. It is that the organization cannot reliably answer what a record refers to, what a field means, what happened before, or who is allowed to see it.

Those four questions are the whole game. A model sitting on top of an organization that can answer them is transformative. A model sitting on top of one that cannot will produce fluent, confident output that two departments will then disagree about, which is worse than nothing, because it looks like an answer.

The four problems, in the order they bite

1. Identity: is this the same thing?

One customer in three systems, spelled three ways. One student who transferred, changed their name, and now exists twice. One supplier that is also a customer under a slightly different legal name.

This is the least glamorous problem in enterprise software and the one that most reliably stops everything downstream. It is also concentrated exactly where you care most. The edge cases in your identity data are rarely the dormant records. They are the transfers, the returning clients, the multi-entity relationships, the people whose situation is complicated. Which is to say: the population any useful model was going to be asked about.

We described the same problem in higher education, where joining advising, enrolment and aid records is the work that has to precede any of the interesting applications. It is the same problem in claims, in manufacturing and in clinical practice. It is always this problem.

2. Definitions: do two people mean the same thing?

Ask two departments what counts as an open case, an active client or a completed order, and you will frequently get two answers, both defensible and both currently in use in a report someone trusts.

Automating on top of that disagreement is the most expensive mistake in this category, because it industrialises it. The output is consistent, fast and confident, and it is wrong for whichever department did not write the specification. We have watched an eighteen-month data platform arrive at a question that turned out to be unanswerable for exactly this reason.

The fix is not technical and it is not expensive. It is a meeting nobody wants to chair.

3. History: can you see what happened?

Models learn from examples of correct answers. A surprising number of operations have never recorded one.

The claims operation we built a platform for tracked 93 fields per case in a spreadsheet, which sounds like a lot of data until you ask what the right answer was for any given field, and discover that the workbook holds the current state and not the decisions that produced it. Who changed the reserve, when, and why is not in there. It was in an email.

If nobody ever wrote down what the correct outcome was, you do not have a training-data problem you can solve with effort. You have a collection problem you can only solve with time, and it should start now rather than at the beginning of the project that needs it.

4. Access: who is allowed to know this?

The last one is the one that surfaces late and kills schedules. A model that can answer any question about your operation is, functionally, a user with permission to read everything.

If your permission model lives in the interface rather than in the data, which is the normal arrangement in older systems, then a retrieval layer over that data has quietly bypassed years of access control. This is discovered in security review, in month five, and it is not a small remediation.

Why "we will clean the data later" always loses

Because the model does not fail loudly on bad data. It fails plausibly.

A pipeline with a broken join does not crash. It produces answers for the records that matched and silently omits the ones that did not, and the output looks entirely reasonable: a report with numbers in it, slightly wrong, in a direction nobody can see. By the time someone notices, the number has been in three management packs.

This is the same argument as the cost of being wrong: what matters is not the error rate but whether the error is caught. Bad data plus a model is the worst case on that axis, because the model's fluency actively disguises the defect.

You do not need clean data. You need this data.

Here is the part that saves projects, and it cuts against how this work is usually sold.

"Fix the data first" is advice that produces paralysis, a two-year programme, and a governance committee. Most organizations cannot fix all their data and should not try, because most of their data is not load-bearing for any decision anyone is waiting on.

The productive version is narrow. Pick one decision somebody cannot currently make quickly. Trace backwards to the fields that decision depends on, usually between five and fifteen of them, not several hundred. Fix identity, definitions, history and access for those fields only. Build the thing. Ship it.

Then do it again for the next decision, reusing the identity work you already did, because that part compounds and almost nothing else does.

This is the same sequencing argument we make about nearshoring operations: something real in production in months, each step paying for the next, rather than a platform that is not usable until it is finished.

The diagnostic, in one question

When a client asks us for an AI system, we ask what decision it would change, and who would make it differently.

If the answer is specific, whether that is "we would stop the shipment," "we would call that client this week," or "we would not pay that invoice yet," then trace the data behind that one decision and you have a scoped, fundable project with a visible result.

If the answer is "we would have better visibility," there is no project there yet. There is a conversation to have first, and it will save whatever the platform would have cost.

The organizations that get real value from AI in the next two years will not be the ones that bought the best models. They will be the ones that spent this period able to say, without a week of reconciliation, what is true about their own operation.

That groundwork is what our data and analytics work is, and it is the least exciting and most durable thing we do.

Working on something like this?

We spend the first conversation understanding what you run on today. No pitch, and no obligation to build anything.

Book a call