Salary delta, time-to-productivity, retention risk, integration tax. We argue upskilling wins below 200 employees because the AI work is 70% domain and 30% LLM; external hire wins above 500 because the platform demands specialist depth. Plus: how to spot the 'AI engineer' who is actually a fine-tuning specialist when what you need is a systems engineer.
Every founder we talk to at the 50-person mark has the same tab open: a job advert for an AI engineer they are not sure they need. Every CTO at the 500-person mark has the opposite problem: several people who have taught themselves prompting, no shared platform, and a board asking why the pilots do not connect to anything. The hire-or-upskill question is really a question about what the AI work at your size consists of, and that changes more between 50 and 500 people than most hiring plans admit.
Our position, from building these systems inside operating companies rather than selling the people who might: below 200 employees, upskill an engineer you already have. Above 500, hire the specialist. In between, it depends on whether your bottleneck is process or platform, and there is a third option that beats both more often than either side of the argument likes to admit.
What does the maths look like at each size?
Four costs decide this, and we are not going to put salaries on them: ours would be wrong for your market by the time you read this. What matters is which way each cost moves as the company grows.
- Salary delta. The gap between what you already pay and what a specialist commands. The number everyone argues about, and the smallest of the four below 200 people, because it buys only the teachable part of the work.
- Time-to-productivity. How long before the person ships something that runs in production on your systems. The insider ships in weeks; the outsider spends months learning where the data actually lives.
- Retention risk. Specialists are recruited continuously, and the one who lands alone with nobody able to review the work leaves first. An upskilled insider is more marketable too, but has tenure, relationships and often equity holding them.
- Integration tax. The cost of connecting the AI work to the CRM, the ERP, the identity system and the people who run them. Near zero for the insider who built those integrations; the largest hidden line for the outsider.
Why does upskilling win below 200 people?
Because of what the work is. At 50 people, and still at 200, the AI work is roughly 70% domain and 30% LLM. The 70% is knowing that the invoice-matching exception happens at month end, that the CRM has two fields called 'owner' that mean different things, that the ops lead will not adopt a tool that adds a click, and that the finance director wants the audit trail before the accuracy number. The 30% is prompting, retrieval, evals, model choice and not trusting the output until the golden set says you can.
The 30% can be taught to a competent engineer in months, and the material to teach it with is now abundant. The 70% cannot be hired in, because it is your company, and it takes the outsider longer to learn than it takes the insider to learn the LLM part. That asymmetry is the whole argument, and every one of the four costs points the same way.
You can teach an engineer the 30%. Nobody can hire in the 70%, because the 70% is your company.
Where the upskilled engineer sits matters too. Below 200 people the bottleneck is process, not platform, so they belong close to operations, working the workflows a COO already owns, with an engineering lead reviewing the code. The reporting line should follow the problem.
Why does hiring win above 500?
Because the work changes shape. Somewhere past 500 people the question stops being 'automate this workflow' and becomes 'give every department a way to automate its own'. That is a platform: shared retrieval, an eval harness every team runs against, cost governance so one team's copilot does not eat another's budget, model routing and observability. It is continuous work, not project-shaped, and it needs someone who has built that platform before.
Three conditions have to hold for the hire to succeed. AI is a permanent part of the engineering agenda, if not the product itself. The backlog is continuous rather than a sequence of projects. And there is an engineering team for the specialist to join, with a manager who can actually review the work. The hire who lands alone in operations, with nobody able to tell good from plausible, is the most common failed placement we see at any size.
What happens between 200 and 500?
It depends, and we can say on what. If your bottleneck is still process (ops-heavy, regulated, workflow-rich), you are a large company with a small-company problem, and upskilling still wins; the insider's 70% is worth more than the specialist's 30%. If your bottleneck has become platform (several teams waiting on the same retrieval layer, model bills nobody owns, pilots that cannot connect to anything), you have a 500-person problem early and should hire for it, provided the manager exists.
The honest tell is whether you can write the job description. If the advert reads as a list of workflows, you want an upskiller. If it reads as a list of infrastructure, you want a specialist. If you cannot write it at all, you are not ready to make either decision well, and that is the case for the third option below.
How do you spot a fine-tuning specialist when you need a systems engineer?
The title 'AI engineer' now covers two people who barely overlap. One trains and tunes models: datasets, loss curves, benchmarks, GPU budgets. The other builds the system around a model somebody else trained: adapters, queues, retries, idempotent writes, evals, observability, the on-call rota. You almost always need the second, because the model is a commodity you call over an API and everything that decides whether the project succeeds sits around it.
The interview tells are consistent. Ask what is still running in production from their last role and who runs it now: the specialist describes a model, the systems engineer a system and the team that inherited it. Ask what they would do first with a workflow whose data sits in three systems that disagree: the specialist wants a dataset, the systems engineer wants the schema and its owner. Ask about the worst incident they were paged for: the specialist has not been paged. Ask how they would know the output had degraded a fortnight after launch: if the answer is not a golden set and a dashboard somebody looks at, keep interviewing.
Fine-tuning has its place: when AI is your product, when you hold proprietary data the base models demonstrably fail on, and only after prompting and retrieval have been exhausted. That describes very few companies below 500 people, and none of them on their first project.
What to do next
Start with the work, not the role. A 20 or 45-minute call with us produces a written view of which workflows are worth automating first and what shape of engineer they need. Where the scope needs defining, a four-day Spec from €5,000 produces a vendor-neutral brief you own outright, which doubles as the job description if you hire and the curriculum if you upskill.
Then consider the third option. A fixed-scope build of the first system, shipped in an average of eight weeks, with the engineer you intend to upskill in the room from week one: reviewing the pull requests, running the evals, owning the integrations. At the end you have a system in production and an engineer who has done the 30% once, with a mentor whose job was to hand it over. Below 200 people that is usually the whole answer. Above 500 it is how you find out what the specialist hire needs to be before you spend months making it.