Engineering OKRs and KPIs: the metrics that actually matter
Most engineering dashboards measure activity: tickets closed, lines of code, hours
logged. None of it tells you whether the team is actually shipping, or getting better at
it. Here are the metrics that do, and how I set them with a team instead of handing them
down.
Engineering OKRs and KPIs are only useful if they measure outcomes a business cares about,
delivery, quality, and predictability, rather than whatever activity is easiest to count.
Get the metric wrong and you get exactly the behaviour it rewards: more commits, not more
value; more tickets closed, not fewer production fires. Get it right, and the team can
tell you, honestly, whether it is getting better.
The metrics that mislead
Lines of code, or commits per engineer. Rewards volume, not judgment. The best fix to a problem is often a deletion.
Tickets closed, or story points burned. Easy to game by slicing work into smaller, less meaningful tickets.
Hours logged, or "utilisation." Measures presence, not output, and quietly punishes the engineer who thinks before typing.
The metrics that actually matter
Four questions, asked as a trend over time rather than a single fabricated number, cover most of what leadership actually needs to know:
Deployment frequency. How often you actually release, trending toward often and boring rather than rare and terrifying.
Lead time for change. How long an idea takes to reach production once someone starts on it.
Change failure rate. How much of what ships needs a fix, a rollback, or an apology.
Time to restore. How fast the team recovers when something breaks, which matters more than pretending nothing ever will.
These are the same shape of metric the DORA research popularised, and the shape matters
more than any specific figure, because a number without your own team's baseline and trend
behind it is meaningless on its own. I do not import someone else's target. I set the
baseline for your team, then watch where it moves.
OKRs vs KPIs, and why the difference matters
A KPI is a number you watch continuously, release cadence, defect rate, lead time for
change. An OKR is a specific, time-boxed objective you are trying to move this quarter,
usually by improving one of those KPIs on purpose. Confusing the two is how teams end up
with a KPI dashboard nobody acts on, or a quarterly OKR nobody can tell is working because
no KPI is attached to it. KPIs tell you the score. OKRs are the play you are running to
change it.
The anti-patterns that quietly ruin OKRs
Most engineering OKR programmes do not fail loudly. They fail by becoming a compliance exercise nobody believes in, usually one of three ways:
Too many objectives. Ten OKRs is not a set of priorities, it is a to-do list wearing a framework's name. Three, tracked honestly, beats ten filed and ignored.
KPIs turned into individual scorecards. The moment a team metric becomes a lever for one person's performance review, people optimise the number instead of the outcome it was meant to represent.
No named owner. An objective with no lead accountable for it belongs to everyone, which in practice means it belongs to no one, and it quietly slips every quarter.
How I set them with a team, not for one
Metrics set in a workshop with the people closest to the work stick. Metrics delivered as a
slide from leadership get gamed, quietly, within a quarter. I run it the first way: ask
the engineers what is actually slow or fragile, turn the honest answer into a small number
of objectives, and put a named lead's name against each one, reviewed every quarter, not
filed and forgotten. It is the same five-part lens an
engineering audit uses to find the gaps in the first
place, turned into an ongoing measurement instead of a one-time diagnosis.
The same discipline is what turns an AI-native
transition into a measured result instead of a demo: without a baseline and a trend to
check it against, "the team adopted AI" is just an anecdote, not evidence anything
actually improved.
Who this is for, and when it's too early
This tends to matter once a team is beyond roughly five or six engineers, once nobody can
hold the whole delivery picture in their head from memory alone, and once there is a
rhythm worth protecting rather than a single sprint to survive. It fits a CEO who has
inherited a team and wants an honest, ongoing read on whether it is improving, and a VP
Engineering who suspects the dashboard in front of the board does not reflect what
engineers actually experience day to day. It is too early if there is no existing delivery
cadence to measure yet; building the first thing to measure comes before measuring it, and
founders in that position are usually better served starting with
the first build itself.
Org structure and recruiting sit right next to this work, not apart from it. The same
conversation that decides what to measure usually decides who owns each objective, which
leads need mentoring into the role rather than replacing, and where the team's next hire
should actually go. Metrics without a structure to act on them are just a report nobody
reads twice.
Zegal · legaltech scale-up · VP Engineering
The three metrics that told me it was working
At Zegal I set the metrics for a remote team of eighteen the same way I do everywhere:
with the leads, not for them. Release cadence, defect rate, and whether a squad could
ship without me in the room became the numbers that mattered, reviewed every quarter,
owned by name. Watching them move told me the team was actually getting better, not
just busier: production issues fell by half, releases settled into a two-week rhythm,
and three squads ran themselves.