Edinburgh Research Explorer

Higher order methylation features for clustering and prediction in epigenomic studies

Research output: Contribution to journalArticle

Original languageEnglish
Pages (from-to)i405-i412
Number of pages12
JournalBioinformatics
Volume33
Issue number8
Early online date29 Aug 2016
DOIs
StatePublished - 1 Sep 2016

Abstract

Motivation: DNA methylation is an intensely studied epigenetic mark, yet its functional role is incompletely understood. Attempts to quantitatively associate average DNA methylation to gene expression yield poor correlations outside of the well-understood methylation-switch at CpG islands.
Results: Here we use probabilistic machine learning to extract higher order features associated with the methylation profile across a defined region. These features quantitate precisely notions of shape of a methylation profile, capturing spatial correlations in DNA methylation across genomic regions. Using these higher order features across promoter-proximal regions, we are able to construct a powerful machine learning predictor of gene expression, significantly improving upon the predictive power of average DNA methylation levels. Furthermore, we can use higher order features to cluster promoter-proximal regions, showing that five major patterns of methylation occur at promoters across different cell lines, and we provide evidence that methylation beyond CpG islands may be related to regulation of gene expression. Our results support previous reports of a functional role of spatial correlations in methylation patterns, and provide a mean to quantitate such features for downstream analyses.

Download statistics

No data available

ID: 34854313