☆ 4.6 Article

Optimally choosing PWM motif databases and sequence scanning approaches based on ChIP- seq data

BMC BIOINFORMATICS (2015)

Journal

BMC BIOINFORMATICS

Volume 16, Issue -, Pages -

Publisher

BMC

DOI: 10.1186/s12859-015-0573-5

Keywords

Transcription factor binding; Binding site; Sequence motif; Motif database

Funding

Polish National Science Centre [2013/09/B/NZ2/03170]
Polish Ministry of Science and Higher Education [N N519 652740]
Polish National Center for Research and Development [ERA-NET-NEURON/10/2013]

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Background: For many years now, binding preferences of Transcription Factors have been described by so called motifs, usually mathematically defined by position weight matrices or similar models, for the purpose of predicting potential binding sites. However, despite the availability of thousands of motif models in public and commercial databases, a researcher who wants to use them is left with many competing methods of identifying potential binding sites in a genome of interest and there is little published information regarding the optimality of different choices. Thanks to the availability of large number of different motif models as well as a number of experimental datasets describing actual binding of TFs in hundreds of TF-ChIP-seq pairs, we set out to perform a comprehensive analysis of this matter. Results: We focus on the task of identifying potential transcription factor binding sites in the human genome. Firstly, we provide a comprehensive comparison of the coverage and quality of models available in different databases, showing that the public databases have comparable TFs coverage and better motif performance than commercial databases. Secondly, we compare different motif scanners showing that, regardless of the database used, the tools developed by the scientific community outperform the commercial tools. Thirdly, we calculate for each motif a detection threshold optimizing the accuracy of prediction. Finally, we provide an in-depth comparison of different methods of choosing thresholds for allmotifs a priori. Surprisingly, we show that selecting a common false-positive rate gives results that are the least biased by the information content of the motif and therefore most uniformly accurate. Conclusion: We provide a guide for researchers working with transcription factor motifs. It is supplemented with detailed results of the analysis and the benchmark datasets at http://bioputer.mimuw.edu.pl/papers/motifs/.

Optimally choosing PWM motif databases and sequence scanning approaches based on ChIP- seq data

Journal

BMC BIOINFORMATICS

Publisher

BMC

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Optimally choosing PWM motif databases and sequence scanning approaches based on ChIP- seq data

Journal

BMC BIOINFORMATICS

Publisher

BMC

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper