| 摘要: |
| 提出了一种新的测序纠错算法.该算法在对测序数据拼接之前对其进行检查,找出并修正测序序列中的错误.该算法将测序数据映射成欧拉超路,并通过一种称为合并变换的等价变换,通过一系列规则的限制和引导,动态地对欧拉超路进行简化.在此过程中,该算法将错误的边和正确的边对应起来,再通过替换纠错过程消除错误.在对T.tengcongensis(TT)和T.whipplei(TW)两个数据集的测试过程中,这种方法分别找出并修正了86%和83%的错误,而原欧拉序列拼接中的纠错算法对这两组数据集的纠错结果只有71%和53%. |
| 关键词: 纠错 序列拼接 DNA测序 欧拉超路 合并变换 |
| DOI: |
| 分类号: |
| 基金项目: |
|
| An Approach to Correcting DNA Sequencing Error |
|
ZHENG Wei-Min,ZHANG Hua,WANG Xiao-Chuan
|
| Abstract: |
| An error correcting algorithm is presented for detecting and correcting errors in the sequencing data before assembly process. The approach maps the sequencing data to an Euler superpath, and simplifies it dynamically by an equivalent transformation named Merging Transformation. In such a process, the algorithm isolates the right edges and error ones so that error paths are substituted and the corresponding errors in the sequencing data are corrected. In two test sets T.tengcongensis and T.whipplei, the algorithm has detected and corrected 86% and 83% errors on the “corrected” sequences respectively, compared with 71% and 53% errors using the original error correcting algorithm in the Eulerian path approach. |
| Key words: error correction fragment assembly DNA sequencing Eulerian superpath merging transformation |