Every care management organization runs into the same operational truth: a small share of cases consumes a disproportionate share of staff time. The 80/20 isn't always 80/20, but the shape is consistent. A handful of consumers — the medically complex, the behaviorally complicated, the families needing constant coordination — eat up case-manager hours that would otherwise go to fifteen or twenty other people. Field managers know exactly which cases these are. They just usually find out after the consumer's been on the caseload for months.
For a national caregiver-support program, the operational question was concrete: could we predict, at the moment of intake, which consumers would turn out to be high-complexity cases? If yes, the team could pre-allocate care resources, set realistic caseload sizes for individual managers, and avoid the demoralizing pattern of one case quietly consuming half a manager's week.
The Promising Setup
The data foundation looked, on paper, like exactly the right starting point. Every consumer entering the program received a standardized assessment — the kind of structured intake form used widely across home-care and caregiver-support contexts, capturing dozens of indicators across medical, functional, psychosocial, and behavioral domains. If anything was going to predict downstream caseload complexity, surely it was this.
To validate the prediction, we needed a ground-truth measure of how complex each consumer's case actually turned out to be. We surveyed care-team members in the field and asked them to rate each of their consumers on a four-point complexity scale, from "stable" through "very high complexity." About 2,400 responses came back, with rich qualitative comments — "diabetic, amputee, totally dependent for care," "major mental health issues," "multiple hospitalizations, frail" — that grounded the ratings in real operational reality.
The plan was straightforward: use a standardized assessment to predict the field's complexity ratings. If the assessment data carried the signal, we'd build a segmentation tool. If it didn't, that itself would be a useful finding.
What the Data Said
Most variables in the assessment had essentially no relationship to field-rated complexity. Of the variables that did show a statistically detectable relationship, a meaningful share moved in the wrong direction — that is, the indicator pointed one way and complexity went the other. Age was the cleanest exception: older consumers genuinely had higher field-rated complexity, in line with intuition. Beyond that, the assessment was largely silent on the question that mattered.
This wasn't a modeling failure. It was an information failure. The assessment data captured what it was designed to capture — clinical and functional state at intake — and the field's "complexity" rating reflects something different: how much sustained attention a case will demand. Those two things are correlated, but loosely. A medically straightforward consumer with a chaotic family situation can be more complex to manage than a medically severe one with strong support at home. The assessment doesn't see the family situation.
What We Built Anyway
Even with the assessment's limited predictive power on its own, the data was strong enough to support a workable segmentation framework. We built two parallel severity scores — one across the medical and functional indicators, one across the psychosocial and behavioral indicators — using the count of "flag" responses on each domain to assign every consumer to a four-tier severity bucket on each dimension. Crossing the two scores produced a 2×2 quadrant model: lower-needs cases in one corner, very-high-needs cases in the opposite corner, and two single-dimension quadrants in between.
About 68% of consumers landed in a high-needs quadrant on at least one of the two dimensions. That distribution alone was operationally useful: it framed the realistic baseline for caseload planning. The framework gave the field a common vocabulary for triaging consumers at intake, even if the underlying prediction wasn't yet sharp enough to drive individual case assignments.
complexity scores collected
domain (low → very high)
bucket on at least one dimension
The Bigger Lesson
The honest finding wasn't comfortable: the standardized assessment data was insufficient to predict the operational outcome the team most cared about. That happens more often than people admit. A regulatory or industry-standard data-collection instrument is designed to capture something, and when an analytics team starts digging in, the natural assumption is that the instrument captures whatever question is being asked. Often it doesn't. The assessment was built to characterize clinical state, not to predict caseload burden — and those are genuinely different constructs.
When standardized data doesn't predict the outcome you care about, the answer isn't to torture the data. It's to be honest about the gap, build the best framework you can with what you have, and identify what additional information would actually move the prediction.
The recommendation went in two directions. Short term: deploy the 2×2 framework as a triage aid, with the explicit caveat that it identifies populations rather than predicting individual cases. Long term: collect the data the assessment doesn't capture — family dynamics, caregiver capacity, social-determinants signals, narrative qualitative content from intake notes — and build a richer model when that data became available.
The methodology pattern recurs across industries. Standardized assessment data in healthcare, standardized customer surveys in services, standardized firmographic data in B2B — each of these is reliable, comparable, and easy to collect. None of them, on their own, reliably predict the outcomes that operators most want to predict. The honest analytics answer is to say so, deliver what the data does support, and build the case for collecting what's actually missing. That's a more useful conversation than pretending the model works better than it does.
Trying to predict an operational outcome from data that wasn't collected for that purpose? Say hello.