Eureka Moments

How an AI Platform for Engineering Teams Got a New Way to Spot Risk

A developer-productivity platform was tracking everything — commits, reviews, response times, throughput. None of it captured the thing engineering managers worry about most. The fix came from an unexpected place.

Engineering analytics platforms have spent the last decade getting very good at counting things. Pull requests opened. Reviews completed. Cycle times. Time-to-first-response. The dashboards are dense, the trendlines are clean, and a manager looking at them can answer almost any question — except the one that actually keeps them up at night. Which is some version of: how fragile is my team, really?

A team can hit every throughput target while quietly accumulating risk. One senior person doing 80% of reviews. Three engineers who have never touched the part of the codebase the new initiative depends on. A handful of people who, if they left, would take half the institutional context with them. None of that shows up in averages. None of it shows up in totals. The numbers look fine, right up until they don't.

An Idea From a Different Discipline

We brought a developer-productivity platform a proposal: import a concept from outside software entirely, and use it to surface the kind of concentration and fragility that traditional engineering metrics miss. The idea wasn't new in its home field — economists have used it for a century to study how unevenly something is distributed across a population. It just hadn't been applied to engineering team behavior. Once it was, it gave a team a single number that captured how broadly or narrowly an activity was spread across the people doing it. A team where everyone contributes equally looks one way. A team where one person carries everything looks completely different. The difference shows up cleanly, in a way no average ever would.

The most useful contribution to a product isn't always a new model. Sometimes it's pointing out that a question the product hasn't been asking has a clean, well-developed answer waiting in another field.

From Idea to Product Feature

We piloted the approach across more than 200 organizations and several million underlying interactions, building benchmarks calibrated by team size — because what counts as healthy distribution looks different across a five-person team and a fifty-person org. Once the norms held up, the platform's engineering team built the concept directly into the product. What had been an outside idea became a native signal — visible to every customer, integrated into the dashboards engineering leaders were already looking at, and updating in real time as team activity flowed through the system.

SAME HEADLINE NUMBERS · THREE VERY DIFFERENT TEAMS TEAM A Healthy distribution TEAM B Some lean on the team TEAM C One person carries it All three teams ship the same total volume of work each week. Only the third one is one resignation away from a problem.
Throughput averages can't tell these teams apart. A concentration measure can — instantly.

Why It Worked

The reason the signal was useful in production wasn't the math. The math is well-understood, taught in undergraduate stats courses, and easy to compute. The reason it worked was that it answered a question engineering leaders were already asking — who's actually carrying this team, and what happens if they leave? — using data the platform already had. The leap was recognizing that a tool from outside the engineering-metrics canon was the natural fit. Once that mapping was made, the build was straightforward.

200+
Engineering organizations
used to validate the approach
Millions
Of underlying interactions
analyzed in the pilot
Shipped
As a native signal
in the production platform

The Broader Pattern

Most analytics platforms have more data than they have ideas about what to ask of it. The biggest unlocks rarely come from a more sophisticated model on the same question. They come from someone surfacing a question the platform hadn't been asking yet — usually by importing a concept from a field where the question is well-developed. Economics, epidemiology, ecology, operations research, behavioral science: each of these has spent a century or more answering questions that are structurally identical to ones a modern AI-era platform has the data to answer but hasn't gotten around to.

Finding those translations is what good outside thinking actually delivers. The model is the easy part.


Building or running an analytics platform that could use a fresh angle on its own data? Say hello.