Job detail for Principal Data Engineer (RWE)
Use AI to assess how you fit
-
Development of data processes for the automated ongoing generation of patient level data (the “data product”) to be used by various business stakeholders for a variety of purposes (e.g. dashboards, reports, studies).
-
Downstream manipulation of datasets after their onboarding from data vendors/partners from “raw” format as provided into useable data structures that will be used to carry out RWE studies, dashboards & other data outputs
-
Transform heterogeneous raw healthcare datasets into reusable data models supporting observational research and epidemiology studies.
-
Occasional conversion of bespoke / one-off datasets (e.g. biomarkers, mutations) to OMOP format (including an understanding of what can and cannot be converted to OMOP format, e.g. to allow analysis to be carried out on residual data that cannot be converted to OMOP).
-
Build FAIR (Findable, Accessible, Interoperable, Reusable) data pipelines and semantic data engineering frameworks to improve discoverability and of healthcare data assets.
-
Create AI-ready datasets that can support generative AI use cases.
Communication
-
Technical engagement with key stakeholders (e.g. epidemiologist, statisticians, market access/health economists) from outside the RWE programming team to ensure a full and detailed understanding of end-user requirement is created and carefully documented. This includes scoping discussions, business analysis and translation of verbalised end-user needs into actionable data structures
-
Detailed technical engagement with colleagues from within the RWE programming team to build data structures required for the generation of RWE study outputs and data products; also support those team members in creating the study outputs where the data engineer’s skillset can add incremental value
-
Liaison & ongoing interaction with IT department to ensure that raw datasets inbound from data partners are fit for the agreed purposes (as per bullet point 1 above)
-
Liaise, where required, with technical staff employed by analysis software vendors (databricks etc)
Documentation
- Maintain clear documentation of data flows, schemas, pipelines, and processes to facilitate onboarding, troubleshooting and auditing.
Quality, Validation & Support
-
Design and carry out detailed testing (data validation and monitoring) approaches for data structures built by self or other members of team to ensure the accuracy and reliability of the data within the data product
-
Troubleshoot any issues encountered with data loading, extraction and transformation (ETL)
-
Work in collaboration with three other members of the Data Engineering team, taking on workload from others as and when required
Required Skills & Qualifications
Domain Expertise
-
Strong understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts.
-
Ability to assess business requirements and recommend appropriate real-world healthcare datasets for analytical use cases.
-
Deep understanding of healthcare data models and healthcare data ecosystems.
-
Strong expertise in OMOP CDM v5.4 , v6, including extensions.
-
Knowledge of healthcare terminologies and standards such as:
o SNOMED CT
o RxNorm
o ICD-10
o LOINC
o HCPCS/CPT
Data Engineering
-
Strong experience in building scalable ETL/ELT pipelines.
-
Expertise in:
o Databricks
o PySpark
o Spark SQL
o SQL
o Delta Lake
-
Experience working with large-scale healthcare and patient-level datasets.
-
Strong understanding of Semantic Data Engineering principles.
-
Experience building FAIR-compliant data pipelines.
-
Experience with cloud-based data platforms and distributed processing frameworks.
Analytics & Visualization
-
Strong Power BI development and data modelling skills.
-
Ability to create reusable analytical datasets for dashboards and studies.
-
Experience designing AI-ready datasets and analytics data products.
Validation & Quality
-
Experience implementing automated data quality frameworks.
-
Strong data profiling, validation, and monitoring skills.
-
Understanding of healthcare data quality assessment methodologies.
Collaboration & Communication
-
Excellent stakeholder management and communication skills.
-
Ability to translate complex business requirements into technical solutions.
-
Experience working with cross-functional global teams.
Disease Area Knowledge
Exposure to one or more of the following therapeutic areas:
-
Oncology
-
Respiratory
-
Immunology & Inflamation
-
Infectious Diseases
Nice-to-Have Skills
-
Working knowledge of R programming.
-
Experience with sparklyR.
-
Experience developing analytical applications using R Shiny.
-
Knowledge of common observational research methodologies.
-
Familiarity with OHDSI tools.
-
Exposure to Azure Data Platform services.
Veramed is a B Corp accredited company which means that we use the power of business to build a more inclusive and sustainable economy meeting the highest verified standards of social and environmental performance, transparency, and accountability.
As an organisation that has people at the heart of it, Veramed is committed to creating a diverse environment and is proud to be an equal opportunities employer. We foster a working culture where employees have integrity, honesty and respect for one another without regard to race, national origin, religion, gender identity or expression, sexual orientation or disability. All qualified applicants will receive equal consideration for employment.