Skip to content
What I do How it works Recent work About FAQ Book an intro call

Engineering OKRs and KPIs: the metrics that actually matter

Most engineering dashboards measure activity: tickets closed, lines of code, hours logged. None of it tells you whether the team is actually shipping, or getting better at it. Here are the metrics that do, and how I set them with a team instead of handing them down.

Engineering OKRs and KPIs are only useful if they measure outcomes a business cares about, delivery, quality, and predictability, rather than whatever activity is easiest to count. Get the metric wrong and you get exactly the behaviour it rewards: more commits, not more value; more tickets closed, not fewer production fires. Get it right, and the team can tell you, honestly, whether it is getting better.

The metrics that mislead

  • Lines of code, or commits per engineer. Rewards volume, not judgment. The best fix to a problem is often a deletion.
  • Tickets closed, or story points burned. Easy to game by slicing work into smaller, less meaningful tickets.
  • Hours logged, or "utilisation." Measures presence, not output, and quietly punishes the engineer who thinks before typing.

The metrics that actually matter

Four questions, asked as a trend over time rather than a single fabricated number, cover most of what leadership actually needs to know:

  • Deployment frequency. How often you actually release, trending toward often and boring rather than rare and terrifying.
  • Lead time for change. How long an idea takes to reach production once someone starts on it.
  • Change failure rate. How much of what ships needs a fix, a rollback, or an apology.
  • Time to restore. How fast the team recovers when something breaks, which matters more than pretending nothing ever will.

These are the same shape of metric the DORA research popularised, and the shape matters more than any specific figure, because a number without your own team's baseline and trend behind it is meaningless on its own. I do not import someone else's target. I set the baseline for your team, then watch where it moves.

OKRs vs KPIs, and why the difference matters

A KPI is a number you watch continuously, release cadence, defect rate, lead time for change. An OKR is a specific, time-boxed objective you are trying to move this quarter, usually by improving one of those KPIs on purpose. Confusing the two is how teams end up with a KPI dashboard nobody acts on, or a quarterly OKR nobody can tell is working because no KPI is attached to it. KPIs tell you the score. OKRs are the play you are running to change it.

The anti-patterns that quietly ruin OKRs

Most engineering OKR programmes do not fail loudly. They fail by becoming a compliance exercise nobody believes in, usually one of three ways:

  • Too many objectives. Ten OKRs is not a set of priorities, it is a to-do list wearing a framework's name. Three, tracked honestly, beats ten filed and ignored.
  • KPIs turned into individual scorecards. The moment a team metric becomes a lever for one person's performance review, people optimise the number instead of the outcome it was meant to represent.
  • No named owner. An objective with no lead accountable for it belongs to everyone, which in practice means it belongs to no one, and it quietly slips every quarter.

How I set them with a team, not for one

Metrics set in a workshop with the people closest to the work stick. Metrics delivered as a slide from leadership get gamed, quietly, within a quarter. I run it the first way: ask the engineers what is actually slow or fragile, turn the honest answer into a small number of objectives, and put a named lead's name against each one, reviewed every quarter, not filed and forgotten. It is the same five-part lens an engineering audit uses to find the gaps in the first place, turned into an ongoing measurement instead of a one-time diagnosis.

The same discipline is what turns an AI-native transition into a measured result instead of a demo: without a baseline and a trend to check it against, "the team adopted AI" is just an anecdote, not evidence anything actually improved.

Who this is for, and when it's too early

This tends to matter once a team is beyond roughly five or six engineers, once nobody can hold the whole delivery picture in their head from memory alone, and once there is a rhythm worth protecting rather than a single sprint to survive. It fits a CEO who has inherited a team and wants an honest, ongoing read on whether it is improving, and a VP Engineering who suspects the dashboard in front of the board does not reflect what engineers actually experience day to day. It is too early if there is no existing delivery cadence to measure yet; building the first thing to measure comes before measuring it, and founders in that position are usually better served starting with the first build itself.

Org structure and recruiting sit right next to this work, not apart from it. The same conversation that decides what to measure usually decides who owns each objective, which leads need mentoring into the role rather than replacing, and where the team's next hire should actually go. Metrics without a structure to act on them are just a report nobody reads twice.

Zegal · legaltech scale-up · VP Engineering

The three metrics that told me it was working

At Zegal I set the metrics for a remote team of eighteen the same way I do everywhere: with the leads, not for them. Release cadence, defect rate, and whether a squad could ship without me in the room became the numbers that mattered, reviewed every quarter, owned by name. Watching them move told me the team was actually getting better, not just busier: production issues fell by half, releases settled into a two-week rhythm, and three squads ran themselves.

-50%
production issues
2 wks
release cadence
3
squads running without me in the room

Want metrics your team actually trusts?

Tell me what you're tracking today, if anything. If OKRs aren't the gap, I'll say so on the first call.

Book an intro call