How Real-World Evidence Data Analytics Falters in Real Clinics
8 min read
The Disconnect Between Clean API Slides and the Midnight Chart-Scribble
In a representative oncology clinic, a clinical research coordinator sits before two monitors at 7:30 PM, manually reconciling a patient's record. The software vendor promised that their machine learning algorithms would automatically ingest, normalize, and deliver research-ready real-world data directly to the sponsor's dashboard. Instead, the coordinator is staring at a scanned PDF of a pathology report from an outside lab, where the critical biomarker status—a key inclusion criterion—is buried in a handwritten margin note.
The failure here is not one of intellect or software capability. It is a failure of execution at the point of care. In the sales presentation, real-world evidence data analytics is depicted as a frictionless pipeline where electronic health records (EHRs), insurance claims, and registry data flow smoothly into an analytical engine. In production, however, this pipeline is choked by the realities of clinical practice. Doctors do not document to satisfy downstream data scientists; they document to get through their shift and keep their patients alive.
This gap between the idealized slide deck and the messy clinic floor is the defining challenge of modern drug development. While the Latin American RWE solutions market is projected to reach $196.7 million by 2030 from $103.1 million in 2025, and the European market is set to double to $2309.8 million over the same period, these capital flows mask a deeper operational friction. We are currently in a half-finished migration, caught between the legacy era of manual chart abstraction and the promised land of fully automated, standardized clinical data.
The Friction Between Structured Promises and Unstructured Progress Notes
When life sciences companies buy real-world data, they often believe they are purchasing a uniform product. The reality, as noted in recent industry analyses, is that not all data delivers the same level of insight. When organizations mix and match datasets without understanding their structural limitations, they risk overlooking the longitudinal elements that make evidence meaningful. This mismatch is particularly acute in complex therapeutic areas like oncology, where patient journeys are non-linear and highly personalized.
For instance, a sponsor attempting to evaluate a new immunotherapy in Europe might rely on EHR data pulled from several regional hospital systems. Under the European Medicines Agency (EMA) guidelines, particularly within the DARWIN EU network, regulatory submissions require rigorous validation of study endpoints. Yet, when the query runs, the system frequently misses patients who discontinued treatment due to low-grade toxicity. Why? Because the side effects were never coded as formal ICD-10 diagnoses. Instead, they were captured as brief, qualitative remarks in a nursing progress note: "Patient complains of mild fatigue, holding dose for three days."
The Failure of Automated NLP in High-Stakes Regulatory Pathways
To an algorithm searching for structured discontinuation events, this patient appears to be actively on therapy. This is where the hype of machine learning in evidence generation meets the hard ceiling of clinical reality. Machine learning models, while increasingly sophisticated, still struggle with the high-cardinality, contextual nuances of clinical narrative. A model trained on oncology notes from one academic medical center often fails when deployed on the unstructured text of a community clinic in a different country, where local idioms and abbreviations distort the semantic meaning.
EHR data migration is less like copying files from one hard drive to another and more like translating a dialect spoken only in a single mountain village into standard medical terminology. When we rely on automated natural language processing (NLP) to perform this translation without human oversight, we introduce systematic errors. A typo in a lab value, a misplaced negation ("no evidence of disease" read as "evidence of disease"), or an outdated problem list can corrupt the entire dataset. For an FDA submission, where the threshold for data integrity is absolute, these errors are disqualifying.
The Clinical Data Fallacy: If a clinical endpoint was not recorded for the purpose of active patient care, no amount of post-hoc machine learning can reliably reconstruct it for regulatory approval.
The Half-Built Bridge From Screen-Scraped EHRs to Federated Data Models
For years, the clinical research industry relied on brute-force data extraction. Vendors used basic screen-scraping and optical character recognition (OCR) to pull text from EHRs, followed by heavy manual curation by offshore clinical teams. This approach is slow, expensive, and difficult to scale. Today, we are witnessing an uneven transition toward federated data networks and standardized common data models, such as the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM).
In theory, OMOP CDM allows different hospital systems to map their local vocabularies—whether SNOMED-CT, RxNorm, or LOINC—to a unified standard. This enables federated analytics, where a sponsor can run a single query across dozens of hospital databases without moving the raw patient data, respecting strict privacy regulations like GDPR in Europe and HIPAA in the United States. Flatiron Health and Tempus specialize in oncology-specific real-world data, while TriNetX provides a global federated network for clinical research, showing how the market is dividing between deep specialty networks and broad federated access.
However, mapping legacy clinical data to OMOP is not a simple click-to-convert process. It requires thousands of hours of manual semantic mapping by clinical terminologists. In a representative mid-sized hospital network, mapping just the laboratory chemistry results to LOINC codes can take six months, only to find that different labs use different reference ranges and units of measurement, throwing off the analytical baseline. The table below outlines how these two paradigms perform in actual clinical operations.
| Operational Dimension | Legacy Manual/NLP Hybrid Pipelines | Federated OMOP Common Data Models |
|---|---|---|
| Data Harmonization | Highly variable; dependent on manual abstractor consistency and local site templates. | Standardized; local vocabularies mapped to global ontologies (SNOMED-CT, RxNorm). |
| Regulatory Auditability | Difficult; source-document verification requires manual re-entry and physical chart access. | High; queries are reproducible and code-based, though mapping logic must be audited. |
| Implementation Overhead | Low upfront cost; high ongoing variable cost for manual curation and data cleaning. | Extremely high upfront cost; requires dedicated clinical terminologists and database administrators. |
| Scale & Query Speed | Slow; data extraction and cleaning cycles take weeks or months per cohort. | Rapid; queries execute across federated networks in minutes once mapping is complete. |
Where the Standardized Pipeline Actually Holds Up
It would be a mistake, however, to dismiss the progress made in real-world data analytics as mere marketing theater. There are specific clinical and operational scenarios where standardized RWE pipelines perform exceptionally well. When we step away from complex, subjective endpoints like "progression-free survival" and focus on hard, objective clinical events, the technology delivers on its promise.
For high-volume, low-complexity safety surveillance, automated RWD pipelines are highly effective. When the analytical goal is simply to monitor for known, well-coded adverse events—such as acute kidney injury following a specific contrast agent administration—structured EHR data and insurance claims are highly reliable. In these cases, the clinical event is unambiguous, universally coded via ICD-10-CM, and tied to a definitive laboratory result like a serum creatinine spike. The analytical engine does not need to parse unstructured clinical notes to find the truth; the structured database fields contain all the necessary evidence.
Similarly, in disease areas with highly structured registry designs, such as certain cardiovascular cohorts, the data collection is integrated into the clinical workflow. Here, physicians use structured templates that mandate the entry of key metrics like left ventricular ejection fraction (LVEF) and blood pressure. When the input is standardized at the point of care, the downstream analytics function exactly as sold. The system delivers rapid, reliable insights that can support label expansions or post-marketing commitments without the need for manual chart review.
Systemic Safeguards for the Messy Middle of Evidence Generation
If we accept that clinical data will always be somewhat messy, the solution is not to wait for perfect software. Instead, we must implement humble, systematic safeguards that acknowledge our fallibility and the limitations of our tools. We must build processes that assume the data is broken and actively work to verify it.
- Establish dual-extraction validation protocols: Rather than relying solely on automated NLP pipelines or a single manual abstractor, a randomized 10% sample of the cohort should undergo independent double-abstraction. When discrepancies arise—such as a disagreement on the exact date of tumor progression—they must be adjudicated by a clinical panel. This process quantifies the error rate of the automated pipeline, allowing biostatisticians to adjust their models accordingly.
- Implement clinical workflow checklists within the EHR: If a hospital wants to participate in lucrative RWE-driven clinical trials, it must make standardizing key data points easy for the clinical staff. This means designing EHR interfaces that require a single click to record a structured response, rather than forcing a physician to navigate three nested menus or write a paragraph of free-text. The checklist, a simple tool of clinical discipline, remains our best defense against incomplete data.
- Validate the semantic mapping layer continuously: Software updates, changes in laboratory equipment, and local coding practices can quietly break data pipelines. A hospital might switch from one local lab code to another, causing a query to suddenly return null values for a critical biomarker. Regular, automated validation queries must run against the database to detect these drifts before the data is exported for analysis.
Frequently Asked Questions
What happens to our clinical trial endpoints when a hospital site updates its local EHR system mid-study?
An EHR system upgrade or migration frequently breaks the underlying database schemas and semantic mappings. When a site transitions from an older version of a system to a new platform, local codes for laboratory tests, medications, and procedures often change or get remapped. This causes downstream queries to return incomplete or missing data for patients enrolled after the transition. To mitigate this, sponsors must maintain a continuous schema-validation pipeline that flags sudden drops in data density or unexpected null values, requiring immediate manual re-mapping of the newly introduced local codes to the standard OMOP model.
How do we handle missing biomarker data in retrospective oncology RWE cohorts without biasing the study?
Missing data is a fundamental reality of retrospective RWD. If you simply exclude patients with missing biomarker data, you introduce selection bias, as patients who receive comprehensive biomarker testing often differ systematically from those who do not. The correct operational approach is to conduct a targeted manual chart review of a representative sample of the "missing" patients to determine if the test was truly never performed, or if the results were simply scanned as an unsearchable PDF image. If the data is truly missing, advanced statistical techniques like multiple imputation by chained equations (MICE) should only be applied after validating that the missingness is at random, a determination that requires deep clinical context rather than automated statistical assumptions.
The Clinical Verdict: Real-world evidence data analytics is a powerful tool for modern clinical research, but only when we treat it as a clinical discipline rather than a software product. Sponsors must stop buying the illusion of the frictionless data pipeline and start investing in the unglamorous work of data cleaning, manual validation, and clinical site engagement. The path to regulatory-grade evidence runs through the messy reality of the clinic floor, and there are no shortcuts.
Related from this blog
- Can Patient Recruitment AI Platforms Solve the Enrollment Crisis?
- AI drug discovery timelines will stall at clinical trials
- EDC Systems Shift Clinical Trial Costs Directly to Sites
- How eCOA and ePRO mobile apps reform trial data pipelines
- Can Clinical Trial Management Systems Run in Real Time?
Sources
- Latin-America Real World Evidence Solutions Market Size, Share,Trends, Growth Analysis Report, 2030 - MarketsandMarkets — MarketsandMarkets
- 7 Real-World Ways RWE Is Transforming Healthcare - Endpoints News — Endpoints News
- Real World Evidence Solutions For Oncology Market (2026 - 2033) - Grand View Research — Grand View Research
- Europe Real World Evidence Solutions Market Size, Share,Trends, Growth Analysis Report, 2030 - MarketsandMarkets — MarketsandMarkets
- Real-World Evidence Meets Machine Learning: What It Takes to Future-Proof Evidence Generation - MedCity News — MedCity News
- Key Stakeholders' Knowledge, Opinions, and Interests on Real-World Evidence in the Regulatory Process—Results of an EU-Wide Survey - Wiley & Sons — Wiley & Sons