Eureka Moments

Using AI to Build Datasets That Don't Exist Yet

A healthcare data client needed practice-level information for thousands of specialty physicians. The dataset wasn't for sale — because nobody had built it. We did, with an AI pipeline.

Sometimes the data you need doesn't exist. Not in any commercial dataset, not in any public registry, not in any government file. The information is out there — scattered across thousands of websites, professional listings, and unstructured text — but it's never been collected, structured, and verified in one place. The traditional answer is to hire a research team and accept that the project will take months. The new answer is different.

For a healthcare data client, the question was practical: out of a definitive list of several thousand specialty physicians, which ones operated their own private, patient-facing practices — as opposed to working only at hospitals or surgery centers? And for the ones who did, what could we learn about each practice: its size, its location, the procedures it offered, the digital footprint it presented to patients?

No vendor sold this data. The information existed, but it lived on practice websites, in physician profiles, and in directories that varied wildly in structure and completeness. Building it by hand would have taken a research team months. Building it well — with consistent structure and reliable confidence scoring across thousands of records — would have been even harder.

The right question wasn't "where can we buy this dataset?" It was "what does it take to build it ourselves, at scale, with quality we can defend?"

The Result

What we delivered was a structured, enriched practice-level dataset built directly from a thin input list. Each confirmed practice came with verified details — address, phone, website, provider count, and a procedure-level profile indicating which services the practice offered. Each record carried a confidence score reflecting how strong the supporting evidence was, so the client could filter or weight downstream usage accordingly.

FROM A PROVIDER ROSTER TO A PRACTICE-LEVEL DATASET BEFORE — INPUT ROSTER Provider ID Name Specialty — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — Three columns. A list, not a dataset. AI PIPELINE AFTER — ENRICHED PRACTICE DATASET ID Practice Address Phone Web N Prv Procedures Conf — — — — — — — — — — — — — .95 — — — — — — — — — — — — — .90 — — — — — — — — — — — — — .95 — — — — — — — — — — — — — .85 — — — — — — — — — — — — — .95 Verified, structured, scored. A real dataset.
A thin input list goes in. A structured, enriched, confidence-scored practice dataset comes out.

Why This Matters Now

Until very recently, this kind of dataset construction simply wasn't economical at scale. You could buy whatever vendors had decided to sell, and for everything else, you accepted a gap. AI changes the calculus. Bespoke enrichment that would once have required a research team and a quarter is now a pipeline that runs in days. The interesting question isn't whether the technology works — it does — it's what your team can do when datasets that didn't exist last year become datasets you can build this week.

The constraint used to be: which datasets are available? The new constraint is: which datasets do you actually need?

Where the Craft Lives

Building these datasets reliably isn't a matter of pointing an AI at the web and accepting whatever comes back. The work is in the structure: defining what counts and what doesn't, building exclusion rules that hold up, sequencing searches in the order that yields highest recall, scoring confidence honestly, validating the output, and engineering the pipeline so it can rerun cleanly as the underlying source data shifts. Done right, the result is a dataset you can stand behind. Done casually, it's a stream of plausible-looking nonsense. The difference is the engineering, not the model.

The same approach applies broadly. Specialty professional rosters. Commercial-real-estate inventories. Long-tail product catalogs. Anywhere the information is real but not yet collected, AI-driven enrichment turns "we wish we had this" into a deliverable.


Have a dataset gap that's been blocking a project? Say hello.