
Non-Coding Driver Mutations
Investigating the 98% non-coding genome to find regulatory mutations that drive breast cancer progression, using more than 5,000 whole genomes from TCGA and ICGC.
Research
I build the computational tools to read the 98% of the genome that doesn't code for proteins, hunting the non-coding driver mutations behind breast cancer.
Research focus
Where the epigenome, machine learning and a little curiosity meet cancer medicine.

Investigating the 98% non-coding genome to find regulatory mutations that drive breast cancer progression, using more than 5,000 whole genomes from TCGA and ICGC.

Developing CNNs, GNNs, MLPs and Hidden Markov Models to predict chromatin accessibility and prioritise candidate driver mutations from multi-omics data.

Building breast tissue-specific regulatory maps by integrating single-cell ATAC-seq, spatial transcriptomics and ENCODE regulatory elements.

Bringing machine learning into the operating theatre: Arduino and ESP32 prototyping, and robot-assisted endomicroscopy at the Hamlyn Centre at Imperial, keeping a probe in focus so a surgeon can read cells at the tumour margin.
Why this work matters
Almost every common cancer-risk variant lives in the 98% we used to call junk. If we can finally read it, we can find the mutations that rewire a healthy cell into a tumour, and the vulnerabilities that could treat it. That's the bet my whole PhD is built on.
Publications

2026
Submitted to npj Genomic Medicine


Let's talk
Collaborations, questions, or just curious about the non-coding genome? My inbox is open.