L

Lawrence Du

About

Detail

Senior AI/ML Engineer
Noblesville, Indiana, United States

Timeline


work
Job
school
Education

Résumé


Jobs verified_user 0% verified
  • GSK
    Senior AI/ML Engineer
    GSK
    Oct 2024 - Current (1 year 10 months)
    • Architected an internal platform for end-to-end fine-tuning of sequence foundation models using PyTorch, PyTorch Lightning, and JAX/Flax variants, enabling multi-team adaptation of RNA and scRNA models with reproducible training templates. • Designed scalable pipelines to fine-tune large sequence models for perturbation prediction from single- cell RNA-seq datasets, integrating distributed training on Google Cloud with Vertex Al Training, TPU v5e, and Ray for parallel hyperparameter sweeps. • Deployed and managed multiple internal sequence model inference services on Google Cloud Vertex Al Endpoints, supporting batch and online inference with autoscaling, model versioning, and A/B rollout strategies. • Implemented highly optimized feature
  • Freenome
    Senior Machine Learning Research Engineer
    Freenome
    Sep 2022 - Jun 2024 (1 year 10 months)
    • Led development of a distributed deep learning platform using PyTorch, Ray Train, Kubernetes, and GCP for large-scale cancer early detection models, enabling >10× speedups via distributed data-parallel training. • Engineered an org-wide MLflow system using Terraform and Pulumi, implementing automated experiment tracking, artifact storage, model lineage, and cloud-orchestrated evaluation jobs across multiple research teams. • Built and productionized multitask learning pipelines integrating methylated cfDNA and proteomic features using PyTorch, improving colorectal cancer risk prediction across a 27,000-patient clinical cohort. • Designed automated model-evaluation workflows using Kubeflow Pipelines and GCP Batch, supporting repeatable cro
  • D
    Owner
    Du Games
    May 2020 - Current (6 years 3 months)
    • Designed and solo-developed Rogue Stargun, a fully shipped VR space-combat simulation built in Unity (URP) with C#, Burst Compiler, and the Job System/DOTS to optimize physics-heavy gameplay on mobile hardware (Meta Quest 2). • Built a hybrid behavior tree + state machine Al system capable of controlling dozens of physics-driven enemy ships simultaneously at >72 FPS on constrained mobile GPUs. • Created custom GPU-instanced rendering pipelines, shader-based VFX, and static batching systems to minimize draw calls; achieved performance targets with aggressive culling, occlusion strategies, and hand-tuned shader graphs. • Integrated text-to-speech synthesis for rapid content generation using APIs from ElevenLabs, Google Cloud TTS, and Azure
  • 23andMe
    Software Engineer - Machine Learning Engineer
    23andMe
    Apr 2020 - Aug 2022 (2 years 5 months)
    • Built a massively parallel statistics computation pipeline (~10 trillion datapoints processed) using AWS Batch, AWS Step Functions, Metaflow, and Athena, reducing genotype-imputation QC runtime from days to hours. • Implemented scalable QC metrics for genotype imputation accuracy, including INFO-score estimation, allele-frequency-dependent QC, and cross-population robustness checks across millions of samples. • Designed and maintained CI/CD infrastructure using Jenkins, Drone CI, Terraform, and containerized build pipelines supporting Python, R, and C++ genome-computation workloads. • Refactored and containerized the Recent Ancestor Locations model into a scalable microservice on AWS using Flask, Gunicorn, S3, and MLflow, supporting high-
  • 23andMe
    Data Scientist
    23andMe
    Nov 2018 - Apr 2020 (1 year 6 months)
    • Primary developer of the modern Recent Ancestor Locations (RAL) algorithm, deployed to >10 million customers worldwide; implemented pipeline components in Python, NumPy, SciPy, and AWS services. • Designed and trained ancestry country-matching classifiers using scikit-learn, xgboost, and probabilistic modeling across millions of customers' genotype datasets. • Built ETL libraries using ORM frameworks (SQLAlchemy) to unify loading of genomic relationships, IBD segments, and reference panel data from cloud storage. • Conducted population structure research using PCA, t-SNE, hierarchical clustering, and novel IBD-based community detection methods. • Implemented tools for population-genetics researchers to access relationship and similarity d
  • T
    Machine Learning Engineer (Bioinformatics)
    The Scripps Research Institute
    May 2018 - Oct 2018 (6 months)
    • Engineered end-to-end RNA-seq and Nanopore analysis pipelines using Common Workflow Language (CWL), Snakemake-style DAG patterns, Conda, and Docker, improving reproducibility and standardizing execution across heterogeneous HPC and cloud environments. Built automated Oxford Nanopore preprocessing workflows supporting basecalling, demultiplexing, and adapter trimming using tools available in 2018 (e.g., Albacore/MinKNOW). • Designed scalable real-time metagenomic diagnostics pipelines integrating Centrifuge, Kraken, and RGI- CARD, enabling sub-hour pathogen detection for clinical microbiology use cases. • Constructed hybrid assembly workflows combining Nanopore long reads and Illumina short reads using Canu, SPAdes, and Pilon, achieving hi
  • J
    Machine Learning Consultant
    Juno Diagnostics
    Sep 2017 - Feb 2018 (6 months)
    • Developed convolutional and fully connected neural networks in TensorFlow 1.x + Keras to detect chromosomal abnormalities from high-throughput NIPT sequencing data. • Designed and implemented a simulation engine using NumPy, SciPy, scikit-learn, and Matplotlib to statistically model read distributions, GC bias, and fetal fraction noise for synthetic training data generation. • Evaluated GPU performance across AWS EC2 GPU instances, local CUDA workstations, and on-prem compute clusters to recommend cost-optimal training infrastructure for deep learning workloads. • Constructed efficient FASTQ-to-feature ETL preprocessing pipelines using Python, pysam, and Pandas, significantly reducing training data preparation time. • Advised founders on
  • Insight Data Science
    Data Science Fellow
    Insight Data Science
    Jan 2017 - Apr 2017 (4 months)
    • Developed a Deep Convolutional GAN (DCGAN) in TensorFlow 1.x to synthesize stylistic cartoon artwork, exploring latent-space disentanglement and deep generative modeling techniques. Built DeepPixelMonster.com, a full-stack web app using Python Flask, Gunicorn, Nginx, and a TensorFlow inference backend to generate artwork in real time on CPU/GPU. • Designed REST APIs for live model inference, including image generation endpoints, caching layers, and workload throttling for cost-efficient deployment. • Implemented automated training pipelines using Python, NumPy, and AWS EC2 instances, including snapshotting and checkpoint management. • Performed hyperparameter sweeps involving convolutional depth, batch normalization configurations, optimi
Education verified_user 0% verified
  • University of California, San Diego
    Ph.D - Biological Sciences, Computational Genomics & Machine Learning
    University of California, San Diego
    Aug 2010 - May 2017 (6 years 10 months)
    Built one of the early applications of convolutional neural networks (CNNs) for cis-regulatory sequence classification using Python, NumPy, and TensorFlow 1.x. • Implemented Deep Taylor Decomposition (Layer-wise Relevance Propagation variant) in TensorFlow, enabling interpretability of transcriptional start site determinants beyond PWM-level motifs. • Developed a GPU-accelerated deep learning training environment using AWS EC2 GPU instances (Kepler/Maxwell), Google Cloud, and custom-built CUDA workstations. • Created MATLAB and Python pipelines for image segmentation, smFISH quantification, and high- dimensional fluorescence analysis using custom-written morphological filters. • Developed statistical models for cis-regulatory activity using
  • Cornell University
    B.A - Biological Sciences, Genetics & Development, magna cum laude
    Cornell University
    Jan 2006 - Jan 2010 (4 years 1 month)