
Non-Coding Driver Mutations
Investigating the 98% non-coding genome to find regulatory mutations that drive breast cancer progression, using more than 5,000 whole genomes from TCGA and ICGC.
Research
I build the computational tools to read the 98% of the genome that doesn't code for proteins, hunting the non-coding driver mutations behind breast cancer.
Research focus
Where the epigenome, machine learning and a little curiosity meet cancer medicine.

Investigating the 98% non-coding genome to find regulatory mutations that drive breast cancer progression, using more than 5,000 whole genomes from TCGA and ICGC.

Developing CNNs, GNNs, MLPs and Hidden Markov Models to predict chromatin accessibility and prioritise candidate driver mutations from multi-omics data.

Building breast tissue-specific regulatory maps by integrating single-cell ATAC-seq, spatial transcriptomics and ENCODE regulatory elements.

Exploring Arduino and ESP32 prototyping as a practical bridge from machine learning research into surgical robotics and autonomy workflows.
Currently exploringWhy this work matters
Almost every common cancer-risk variant lives in the 98% we used to call junk. If we can finally read it, we can find the mutations that rewire a healthy cell into a tumour, and the vulnerabilities that could treat it. That's the bet my whole PhD is built on.
Publications

2026
Submitted to npj Genomic Medicine


Let's talk
Collaborations, questions, or just curious about the non-coding genome? My inbox is open.