Dev productivity is not just PR

Is a PR (pull request) a good way to measure developer productivity? yes and no.

It gives us a visible count, but the count is not the work. Measuring a developer only by pull requests is like measuring a farmer only by how many crops they harvest in one hour. The number is real, but it says little without knowing the crop, the season, the preparation, and the quality of the result.

A pull request can help us ask better questions. Did activity suddenly drop? Are changes waiting for review? Did the team split work into smaller pieces? Is one person stuck on an issue that needs help?

Those are useful questions. "Who produced the most?" is usually not.

Pull requests also hide differences in the work. One change may be a small config update. Another may follow days of debugging, design, or coordination. A senior engineer may submit fewer PRs while reviewing other work, planning an migration, or removing blocker for other people. Ten failed experiments and one correct fix still appear as one final change in most of situation.

This is why I would treat PR count as an signal, not an individual score. If it becomes a goal, people can cheat e.g, increase the count by splitting work differently without improving anything for the product.

How about the delivery system with DORA metrics

Leadership often sees a dashboard without seeing the team's daily work. A tech lead or engineering manager has to connect those numbers to context. What changed during this period? Which constraint is slowing the team? Where should we look more closely?

DORA's current software delivery guidance is a reasonable place to start. Its 5 metrics cover change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. These are signals about a team's delivery system. But they are not a ranking system for individual developers, and they should not turn directly into actions without investigation, in another term - the team signal

SPACE vs DORA

DORA focuses on how safely and quickly software moves through delivery. SPACE is broader. The SPACE framework says developer productivity cannot be reduced to one dimension. It considers satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow.

The 2 frameworks are not opponents. I can use DORA to see a delivery trend, then use SPACE to avoid explaining that trend with activity alone.

For pull requests, I might ask:

  • Are developers satisfied with the review process?
  • Do merged changes produce the intended result?
  • How much review and change activity is happening?
  • Can people get useful feedback from the right teammates?
  • How long does a change wait, and how often is focused work interrupted?

I would not measure every possible item. I would choose a small set based on the problem, then keep a balancing signal nearby. If review speed improves, what happens to escaped defects or rework? If activity rises, what happens to satisfaction and flow?

How to make important invisible work discussable

The same approach helps when asking for investment in developer experience or a platform team. Start with the real problem, not the tool you want to buy or build. Draft the story. Find where people lose time, dig into why, and collect enough data to test that explanation.

The goal is not to make every useful action countable. It is to make important constraints visible enough that the team can decide what to change.

AI nowaday can improve the count, hurt the system and potentially cheat the productivity game

AI makes the distinction even more important. I use it to check my understanding, surface edge cases I missed, or stay in flow when a small question would otherwise stop me. Those are practical gains.

But AI can also generate more code to review, introduce funny bugs, and encourage developer to accept an answer they cannot yet evaluate. Over time, relying on it carelessly may weaken skills and make the output quality go down.

So I would not ask only whether AI increased pull requests or lines changed. Did it shorten the full delivery cycle? Did review effort or rework increase? Did the developer understand the change? Did it protect flow or merely move the interruption to someone else?

Productivity metrics are valuable when they lead to questions like these, but from management view it's not easy to use put these insights in a report so they usually choose metrics/output to evaluate team / ic instead of treating them as signal to understand what the real insights.