AI-powered data prep for real-world auditability

bfPREP™

Data Harmonization

Analysis-ready data, proven on a trial the industry had shelved. Harmonize messy clinical and biological data into an auditable foundation, so every downstream analysis runs on firm ground.

0 %

of organizations lack or are uncertain about the right practices.¹

Data Management Practices for AI

Clinical trial records and other valuable data can be trapped in thousands of PDF reports. Multi-site studies that theoretically measure the same endpoints may inadvertently use different instruments, measurement ranges, or lab-derived protocol differences. Genomic datasets could have bias and conflicting annotations, and so much more.

Combining these typical sources of information without proper biological expertise doesn’t produce actionable insights; it just produces data clutter and noise.

The Input

What You Provide

Your team provides data as it is. No pre-cleaning and no standard formatting required.

Clinical trial reports, PDFs, scanned clinical annotations, and hand-written notes

Multi-site studies, clinical performance studies with mismatched endpoints and disparate study size

Multi-omic data sets derived from different technologies and annotation versions

Data records fragmented across disconnected data systems, spreadsheets, and computer drives

The Work

What We Do

bfPREP™ automates data harmonization and ensures trackability of critical human decisions, necessary for clinical data compliance, reviews, and audits.

01

Extract

Pull raw data out of your existing databases and file it into one place

02

Clean

Fix errors, fill gaps, and remove duplicates so the data is reliable and easy to understand

03

Standardize

Put everything into a consistent format across different datasets

04

Supplement

Enrich the data with publicly available or derivable information for additional context

05

Link

Connect the dots across data types (genomics, drug response, pathology, etc.), so subsequent causal network models link to all available data in a single record

Every connection, mapping, and feature is a logged, validated, scientific judgement call performed by a scientist. Not a guess by a black-boxed system.

Experts in the Loop

Human oversight is embedded at the checkpoints that matter most.

Your scientific team signs off on the decisions that govern every finding before analysis can even begin.

Mappings reviewed

Proposed data mappings and category schemes confirmed before use

Features validated

Clinical relevance of engineered features verified by scientists

Extractions resolved

Ambiguous extractions analyzed and settled, not guessed

Distributions checked

Statistical distributions verified against source documentation

The Output

What You Receive

Cleaned, harmonized data that is traceable and auditable. You understand the reasoning and give approval before implementation on your precious datasets.

Harmonized, analysis-ready datasets

Linkages on firm biological evidence, flagged separately from judgement calls

A complete record of how they were prepared

Residual uncertainty made explicitly transparent, not hidden

The Outcome

What it Supports

A foundation precise enough to trust downstream, instead of every later step in your AI journey struggling to compensate for upstream data problems.

Completed Engagement – Eleison Pharmaceuticals & H. Lee Moffitt Cancer Center 

A Failed Phase III Trial, Re-analyzed.

Glufosfamide in advanced pancreatic cancer was negative at the population level and headed for the shelf.
~ 0

pages extracted, structured, and validated

0 weeks

from fragmented files to analysis-ready data

0
reproducible patient subgroups discovered
0 x

greater mean survival in Cluster A identified upon re-analysis

bfPREP™ harmonized the data; bfLEAP® identified the subgroups; Moffitt oncologists confirmed them before submission. Cluster A showed 149.3 vs. 58.2 days survival (glufosfamide vs. best supportive care, p=0.065) and the lowest baseline glucose, consistent with the drug’s glucose mediated mechanism. Presented as a poster at ASCO GI 2026.²

Feasibility Assessment

Request a feasibility assessment with our data or your own to explore what’s possible.

Reference 1: GARTNER, 2025. 2. Naleid, N., Gosik, K., Tambe, A., Savkli, C., Beltran, J. F., & Kim, R. D. (2026). Data-driven subtyping and differential glufosfamide benefit in pancreatic adenocarcinoma. Journal of Clinical Oncology, 44(2_suppl), 755. https://doi.org/10.1200/JCO.2026.44.2_suppl.755