An honest read on what is working in cost-intelligence for QS firms, contractors, and procurement teams (comparable cost matching, risk-flag generation, document QA) and what is still vapourware. Plus the data-shape and currency challenges that determine whether your firm can attempt this at all.

Every quantity surveyor we have spoken to this year has seen a demo in which a model reads a bill of quantities and produces a cost plan. Most have also noticed that it ran on the vendor's data, not theirs. The gap between those two facts is this brief.

We build cost-intelligence layers on spec for QS firms, contractors and procurement teams, and we do not sell a product, so we have no reason to tell you the tooling is further along than it is. Some of it works now and is quietly paying for itself. Some of it is a slide. And whether your firm can attempt any of it is decided before a model is chosen, by the shape of your cost data and how you have handled currency.

What is AI cost intelligence actually doing for QS firms today?

Three use cases have moved from pilot to daily use in the firms we work with. They share one property: the model is never the source of a number. It finds, flags and checks, and a surveyor still signs. None of them predicts a price; all of them make the surveyor faster at the part of the job that was never the skilled part.

  • Comparable cost matching. Given a line item, find the closest priced items in the firm's own history, across projects, countries and years, and show them side by side with the adjustments applied. What makes it work: a canonical unit, currency and period on every historical line, so that closest means the same thing everywhere. Without that, closest is a string match, not a comparison.
  • Risk-flag generation. Read a tender return, a subcontract package or a cost plan and raise the lines that deserve a second look: a rate well outside the firm's own range, a quantity that disagrees with the schedule, a provisional sum sitting where a measured item should be. What makes it work: a reference range built from the firm's cleaned history, and a flag threshold set by the cost team rather than the vendor.
  • Document QA. Check the documents a QS produces and receives against each other before a cost report goes out. Does the bill match the specification, does the tender sum match the priced schedule, are the exclusions in the quote also in the comparison? What makes it work: the checks are written as rules by the people who do the job, with the model doing the reading and reconciliation at volume.

What is still vapourware?

The honest list is longer than the working one, and it is where most of the marketing budget sits. If a pitch leans on any of these, ask to see it run on your data.

  • Fully automated cost planning from drawings or a BIM model. Takeoff quality varies by discipline and by who produced the model. A pricing layer on an unreliable quantity is a confident wrong number.
  • Market-rate prediction. A model that tells you what concrete will cost next quarter is extrapolating from a history that is thin, biased towards the deals that closed, and priced in currencies that moved. It is a forecast wearing a decimal point.
  • Upload your spreadsheets and it just works. Any tool that promises intelligence on data it has not been made to clean is describing your data as it wishes it were. The cleanup is the project; the model is a feature.
  • Cross-firm benchmarking pools. The idea is sound. The data is not comparable between firms, often not within one, and nobody has solved consent and confidentiality at the line-item level.
  • An agent that negotiates with suppliers. We have not seen one a procurement lead would let near a contract, and we would not build one.

Why do data shape and currency decide whether you can start?

Everything on the working list rests on one condition: two line items from different programmes can be put next to each other and the comparison defended. It has two halves.

The first half is shape. A firm's cost history lives in bills, cost plans, tender returns and final accounts, written to be read, not queried. Concrete is recorded per cubic metre in one office and per bag in another; a door is a leaf here, a set there, an aperture somewhere else; the same supplier appears under several spellings; some rates are quoted, some contracted, some paid. Every one of the seven dirty-data problems we have written about before shows up in cost data. Until each line carries a canonical unit, supplier and date, a model can only match strings, and matched strings produce plausible comparisons that are wrong in ways nobody notices until the margin does.

The second half is currency, and it is the one firms underestimate. Two rates in the same currency a few years apart are not comparable even before conversion; inflation, exchange movement and supplier currency clauses all compound. The only defensible approach is conversion to a single base currency at the date of the transaction, with the rate used stored beside the line, and no back-conversion afterwards. Tax handling rides on the same problem. One country records gross, another net, and a model that treats them alike will conclude the net country is cheaper.

If your history does not carry those fields, you cannot start with the model. You can start with the cleanup, which is where the value sits.

In cost intelligence the model is the cheap part. The data that lets two line items be compared and defended is the whole cost, and the whole value.

What did it take on a real global programme?

A Danish quantity-surveying firm running global construction programmes came to us with this exact problem. Concrete, doors, glazing and fixtures, sourced in different markets at different times against different supplier conventions, and a cost team that wanted to compare a door in one city with a door in another and defend the number.

We did the data layer first. Over a six-week cleanup phase we normalised units, currencies, periods and supplier conventions across 6 countries, so the same line item read the same wherever it appeared and 100% of the firm's cost data sat in one shape. Only then did we build the pricing intelligence layer on top: cost models and benchmarks the team leans on for live programmes.

Two things about that sequence matter. The cleanup was the larger piece of work and the model the smaller, the reverse of how every vendor prices it. And the layer is trusted because any benchmark traces back to real lines with the adjustments visible. The intelligence is the data, made comparable.

What to do next

Locate your firm honestly. Pick one line item you price often and ask whether your own history can return every instance of it, across offices and years, in one unit and one currency, without an analyst rebuilding it by hand. If it can, the working use cases are within reach. If it cannot, the first project is the cleanup, scoped to the decision it should support, and anyone selling you a model before that is selling you the vapourware list.

That is the shape of a J Labs engagement. A 20 or 45-minute discovery call to place your data and name the person who can say what canonical means for it. Where it makes sense, a four-day Spec from €5,000 that puts the cleanup and the first use case in one written brief you own outright: which dirty-data problems apply, in what order, and what the layer on top will and will not answer. Then a fixed-scope, fixed-price build; the average ships in eight weeks. No product, no platform, and no number a surveyor cannot defend.

Or skip ahead and talk through it directly