TempoMAGE: A deep learning framework that exploits the causal dependency between time-series data to predict histone marks in open chromatin regions at time-points with missing ChIP-seq datasets

dc.contributor.authorHallal, Mohamad M.
dc.contributor.authorAwad, Mariette
dc.contributor.authorKhoueiry, Pierre H.
dc.contributor.departmentBiochemistry and Molecular Genetics
dc.contributor.departmentBiomedical Engineering Program
dc.contributor.departmentDepartment of Electrical and Computer Engineering
dc.contributor.departmentSpecialized Clinical Programs and Services
dc.contributor.departmentPillar Genomics Institute of Precision Medicine
dc.contributor.facultyFaculty of Medicine (FM)
dc.contributor.facultyMaroun Semaan Faculty of Engineering and Architecture (MSFEA)
dc.contributor.institutionAmerican University of Beirut
dc.date.accessioned2025-01-24T11:38:12Z
dc.date.available2025-01-24T11:38:12Z
dc.date.issued2021
dc.description.abstractMotivation: Identifying histone tail modifications using ChIP-seq is commonly used in time-series experiments in development and disease. These assays, however, cover specific time-points leaving intermediate or early stages with missing information. Although several machine learning methods were developed to predict histone marks, none exploited the dependence that exists in time-series experiments between data generated at specific time-points to extrapolate these findings to time-points where data cannot be generated for lack or scarcity of materials (i.e. early developmental stages). Results: Here, we train a deep learning model named TempoMAGE, to predict the presence or absence of H3K27ac in open chromatin regions by integrating information from sequence, gene expression, chromatin accessibility and the estimated change in H3K27ac state from a reference time-point. We show that adding reference time-point information systematically improves the overall model's performance. In addition, sequence signatures extracted from our method were exclusive to the training dataset indicating that our model learned data-specific features. As an application, TempoMAGE was able to predict the activity of enhancers from pre-validated in-vivo dataset highlighting its ability to be used for functional annotation of putative enhancers. © 2021 The Author(s) 2021. Published by Oxford University Press. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.
dc.identifier.doihttps://doi.org/10.1093/bioinformatics/btab513
dc.identifier.eid2-s2.0-85122196952
dc.identifier.pmid34255822
dc.identifier.urihttp://hdl.handle.net/10938/29009
dc.language.isoen
dc.publisherOxford University Press
dc.relation.ispartofBioinformatics
dc.sourceScopus
dc.subjectChromatin
dc.subjectChromatin immunoprecipitation sequencing
dc.subjectDeep learning
dc.subjectHistone code
dc.subjectRegulatory sequences, nucleic acid
dc.subjectRegulatory sequence
dc.titleTempoMAGE: A deep learning framework that exploits the causal dependency between time-series data to predict histone marks in open chromatin regions at time-points with missing ChIP-seq datasets
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
2021-2295.pdf
Size:
1.41 MB
Format:
Adobe Portable Document Format