Five tiers of AI readiness, from 'spreadsheets and Drive' to 'warehouse + governance', and the use cases that actually work at each tier. Kills the myth that you need a Snowflake bill before you ship anything useful. Honest about the ceiling at each tier, too.
'AI readiness' has become a synonym for 'buy a data platform'. A COO asks what it would take to put AI on the procurement data, and the answer comes back as a reference architecture: warehouse, pipeline tool, catalogue, governance programme, and a data team to run it. Then there is a Snowflake bill and still nothing in production.
We think that answer is backwards. The companies we work with, 50 to 5,000 people with years of data spread across an ERP, a CRM, an old database and a folder of spreadsheets, do not need a data lake to start. They need to know which tier of readiness they are at, which use cases work there, and where the ceiling is, because that is what stops the first project being scoped one tier too high.
Do you need a data warehouse before you start with AI?
No. You need a decision the data should support, someone with the authority to say what clean means for that data, and a use case that lives inside your current tier. That is the whole entry ticket.
The myth persists because the people selling readiness assessments sell the top tier. A reference architecture is easy to draw, easy to price, and it moves the risk to you: if the use case never arrives, the platform still got built. We have watched companies arrive at the use case with the same dirty data they started with, now in a more expensive place. A model prompted on it still lies, just faster.
What are the five readiness tiers?
A company with a warehouse licence and no canonical supplier list is at tier two with an expensive hobby. Place yourself at the highest tier whose description is entirely true, not the highest you have paid for.
- Spreadsheets and a shared drive. What you have: Excel or Sheets, a drive full of documents, and exports pulled from the ERP by hand. What works: question-answering over the documents you already hold, extraction from PDFs and emails into a structured sheet, classification and drafting; anything where one document is the whole input. The ceiling: anything that needs two sources to agree. The answer is only as good as the latest version of the file.
- Operational systems, unconnected. What you have: an ERP and a CRM in daily use, perhaps a database from an earlier era, each with its own names, units, currencies and periods. What works: use cases that live inside one system. Call summaries in the CRM, invoice classification in the ERP, supplier-name resolution, anomaly flags on a single ledger. The ceiling: any cross-system question. The honest answer is a reconciliation spreadsheet, and a model cannot fix what that spreadsheet hides.
- Standardised and connected. What you have: one canonical unit, currency, period and name for everything in scope, decided by someone with the authority to decide, and the systems connected so a record entered once is right everywhere. The reconciliation spreadsheet has retired. What works: pricing intelligence, cost benchmarks, cross-system reporting, classification at volume. A global quantity-surveying firm reached this tier when we standardised its procurement data across 6 countries in a six-week cleanup phase; the pricing intelligence layer went on top. The ceiling: history and maintenance. Everything before the cleanup keeps its old shape, and the standard holds only while someone owns it.
- One source of truth with an owner. What you have: a single governed dataset that the board pack, the dashboards and the finance team all read from, a named data owner, written definitions and scheduled quality checks. What works: forecasting, benchmarks over time, board numbers that match from one meeting to the next, models retrained without a fresh cleanup. The ceiling: volume, variety and self-service. New questions still go through the people who built the source.
- Warehouse and governance. What you have: a warehouse, lineage, access control, a governance programme and, by now, a data team. What works: everything above at scale, plus self-service analysis and machine-learning pipelines with proper monitoring. The ceiling: organisational rather than technical. It costs a team to run, and without a use case that needs it, it is an expensive place to store answers tier three already gave you. Most companies of 50 to 5,000 people never need it.
The warehouse is the top tier of readiness, not the entry ticket. Pick the use case from the tier you are at, and let it pay for the next one.
Which tier are you actually at?
Most companies place themselves one tier above the evidence, because they count what they own rather than what is true. Our test is one question: can the data answer a specific business question without someone rebuilding it by hand? What did we pay for the same item in each country last year, in one currency and one unit?
If that needs an analyst and a rebuild, you are at tier two, whatever the licence says. If it comes back in one currency and one unit with no caveats, tier three. If the number matches the one the board saw last month, tier four. Be honest, because a use case built one tier too high does not fail loudly. A cross-country benchmark on tier-two data produces a plausible number, the team prices the next bid on it, and nobody finds out until the margin does.
How do you move up a tier without a data team?
You do not build the platform. You do the cleanup in the right order, scoped to one dataset and the decision it should support, and you let the first use case justify the next tier.
The order matters more than the tooling. Currency, unit of measure and supplier resolution first, because everything else stands on them. Period alignment and country-specific tax handling next. Supplier hierarchies and the gaps that are not missing at random last, because they need stable foundations.
Three things make this possible. A named owner with the authority to say what canonical means: one unit, one currency, one period, one supplier name. If that person does not exist yet, naming them is the first step. Cleanup where the data already lives: we standardise in the Postgres, the SQL Server, the ERP export and the spreadsheet, then connect them; nothing gets ripped out, and wrap, rebuild or replace is a decision we make with you and show our working on. And artefacts small enough to maintain: a supplier resolution table, a unit dictionary, a documented canonical date. An ops lead can own those. A warehouse, they cannot.
That is the move from tier two to tier three, where most of the value sits for an established business. Tier four follows when the second and third use cases need the same source; tier five when volume demands it, by which point the use cases are paying for the team.
What to do next
Locate yourself honestly. Name the decision the data should support and the person who can say what clean means. Then scope the cleanup and the first use case together, as one piece of work. Scoped apart, the cleanup becomes a platform and the use case becomes a demo.
That is the shape of a J Labs engagement. A 20 or 45-minute discovery call to place the tier and name the owner. Where it makes sense, a four-day Spec from €5,000 that puts the cleanup and the first use case in one written brief you own outright: which dirty-data problems apply to you, in which order, and what the use case on top will and will not answer. From there, a fixed-scope, fixed-price build; the average ships in eight weeks. The tier you are at, the use case that works there, and the cleanup that earns the next one.