Cohort Lab harmonizes EHR and claims data into the OMOP Common Data Model — so your team spends its time on studies, not on cleaning data.
Ask any research data team where the time goes and the answer is rarely the model. It's reconciling source systems, aligning timestamps, chasing coding changes, and arguing about whether two fields mean the same thing.
The OMOP Common Data Model exists to end that argument — one standardized structure and vocabulary, so the same analytical tools run against any conforming dataset without moving sensitive records. It's adopted across more than 30 countries and hundreds of millions of patient records. But getting your data into it is a real engineering project, and it's the part nobody wants to own.
That's the work Cohort Lab does.
Start small and fixed-scope. Expand only if the first piece earns it.
Before anyone writes ETL, you need an honest picture of what's actually in the data.
The ETL itself — built to be read, audited, and maintained by your team, not only by ours.
A CDM isn't a delivery, it's a living asset. Source systems change; the model has to keep up.
OMOP CDM and the OHDSI toolchain as they're actually specified — not a proprietary variant that locks you in. Your CDM should outlive the engagement.
The OHDSI toolchain is open source and so is the work built on it. Every mapping decision is documented and reviewable, because in research the provenance is the result.
Every delivery ships with its data quality profile and known limitations written down. If something is a proxy or a compromise, you'll read it from us first.
Cohort Lab exists to do one thing properly: the observational research data layer. Not analytics dashboards, not another platform to buy into — the unglamorous engineering that turns disparate clinical and claims data into something a study can actually run against.
The practice is built on enterprise data architecture experience in regulated industries, including pharmaceutical R&D — where the systems scientific research depends on carry validation, provenance, and governance constraints that general-purpose data teams rarely encounter. That context is the difference between an ETL that runs and one that holds up to scrutiny.
This is a new practice and we'd rather say so plainly than dress it up. What we bring is deep familiarity with exactly this kind of data, applied to a well-specified open standard — and a preference for fixed-scope first engagements so you can judge the work before committing to more of it.
Both run on synthetic data, both are public, and both state what they don't do. Judging a data practice on a deck is hard; judging it on code is not.
A messy EHR extract converted to OMOP CDM v5.4 — source profiling, a documented mapping specification, and automated data-quality checks enforced by CI on every push. Unmapped codes land at concept_id = 0 and appear in the report rather than being quietly dropped, because a mapping rate that ignores its failures isn't a mapping rate.
Four agents turn a clinical research question into a feasibility-assessed cohort with an attrition table. Tools are served over an MCP server, and the guardrails sit between the agents rather than inside the prompt — schema and vocabulary validation, minimum cell-size suppression enforced inside the tool server, and step and token budgets. Swap the model and the guarantees still hold.
The model never writes SQL and never counts. It proposes a constrained concept set; deterministic code does the arithmetic. A 15-case golden set gates CI, and a second check fails the build if any metric drifts between runs — an evaluation number you cannot reproduce is not a measurement.
The model appears in exactly one plane. Everything above decides what runs; everything below decides what is allowed and what is recorded.
Tell us what data you have and what you're trying to study. No prepared brief required — the messy version is more useful anyway.
A fixed-scope, fixed-price engagement that profiles your source data, drafts the mapping specification, and gives you an honest read on effort, risk, and what won't map cleanly.
ETL pipelines, vocabulary mapping, and quality validation — delivered incrementally so you see a working subset early rather than a big reveal at the end.
Documented, version-controlled, and walked through with your team so they can run and extend it. Ongoing support if you want it, not because you're stuck.
The first conversation is free, and useful whether or not we work together — worst case you leave with a clearer view of the effort involved.
cohortlab@omazesoft.com