Job detail for Research Scientist

G
Research Scientist
George Washington University
Todayvia fourdayweek

Use AI to assess how you fit

The Biostatistics Center (BSC) of the Milken Institute School of Public Health is an off-campus research facility of The George Washington University located in Rockville, Maryland. The BSC serves as the coordinating center for large scale multi-center clinical trials and epidemiological studies funded by federal agencies including the National Institutes of Health. The BSC is a leader in the statistical coordination of major medical research programs of national and international scope. Visit our website at: www.bsc.gwu.edu.

We are seeking a Research Scientist specializing in Machine Learning, Data Science, Data Harmonization, and Synthetic Data Generation to lead the integration, standardization, and privacy-preserving algorithmic modeling of complex, multi-site datasets. In this role, you will bridge the gap between complex data infrastructure, cutting-edge machine learning, and synthetic data generation—building automated transformation pipelines, harmonizing disparate clinical/observational data structures (e.g., OMOP CDM, FHIR), and generating high-fidelity synthetic datasets to accelerate secure collaborative research without compromising data privacy. Experience with high-dimensional multi-omics data analysis and integration is highly desirable.

Key :

Synthetic Data Generation & Privacy Engineering

  • Generative Modeling: Design, train, and validate generative machine learning models (GANs, VAEs, diffusion models, and LLM-based tabular synthesizers) to generate high-fidelity synthetic tabular, longitudinal, and multi-omic datasets.
  • Privacy Assurance & Risk Assessment: Implement rigorous privacy-preserving methodologies (differential privacy, membership inference attack testing, re-identification risk metrics) to guarantee synthetic datasets meet strict governance and compliance standards.
  • Utility & Fidelity Evaluation: Establish automated benchmark suites comparing distributional fidelity, feature correlations, cross-sectional/longitudinal validity, and downstream task performance between real and synthetic cohorts.

Data Harmonization & Infrastructure

  • ETL & Data Standardization: Design, execute, and maintain scalable ETL/ELT pipelines to map complex, longitudinal datasets from multi-center registries and disparate source systems into common data models (e.g., OMOP CDM, PCORnet).
  • Ontology Mapping & Quality Control: Implement vocabulary mappings (SNOMED, RxNorm, LOINC, ICD-10) and robust data quality frameworks to handle missingness, inconsistent formatting, and cross-site measurement variation across clinical and molecular datasets.
  • Registry & Pipeline Integration: Architect automated validation workflows ensuring privacy preservation, site blinding, and reproducible data harmonization across large-scale cohort studies.

Machine Learning, Omics & Statistical Modeling

  • Multi-Omics Data Integration: Analyze, harmonize, and model high-dimensional biological data layers (e.g., genomics, transcriptomics, metabolomics, proteomics) alongside EHR and clinical registry data.
  • Algorithmic Modeling: Develop, evaluate, and deploy advanced machine learning models (gradient boosted trees, neural architectures, ensemble learning) and semi-parametric causal inference techniques on high-dimensional harmonized and synthetic data.
  • Longitudinal & Trajectory Analysis: Build statistical and ML frameworks for analyzing longitudinal trajectories, biological aging markers, and time-varying exposures.
  • Model Explainability & : Conduct rigorous cross-validation, sensitivity analyses, and evaluations to ensure model fairness, robustness, and biological/clinical validity when trained on real vs. synthetic inputs.

Collaborative Research & Leadership

  • ** Collaboration:** Work closely with , , privacy officers, software engineers, and domain experts to translate scientific hypotheses into secure, end-to-end data products.
  • Methodological Dissemination: Lead the drafting of peer-reviewed scientific publications, technical whitepapers, and method documentation regarding privacy metrics, synthetic data utility, and multi-omic modeling.
  • Pipeline Open-Sourcing & Mentorship: Contribute to reproducible codebases (R, Python, SQL) and mentor junior data engineers and graduate research assistants.

Performs other work duties as assigned. The omission of specific duties does not preclude the supervisor from assigning tasks logically related to the position.