Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revision Previous revision
Next revision
Previous revision
bioinformatic_tools_to_detect_microsatellites_loci_from_genomic_data [2011/10/13 16:16]
anniearchambault Update for Fireusat
bioinformatic_tools_to_detect_microsatellites_loci_from_genomic_data [2013/08/08 19:21] (current)
anniearchambault
Line 23: Line 23:
 ^Mreps ​    | 2003     | ?     | [[http://​scholar.google.ca/​scholar?​cites=7808444266556719865&​as_sdt=2005&​sciodt=0,​5&​hl=fr|156]] ​    | Kolpakov et al 2003(([[http://​www.nar.oupjournals.org/​cgi/​doi/​10.1093/​nar/​gkg617|Kolpakov,​ R., Bana, G., and Kucherov, G. (2003). mreps: efficient and flexible detection of tandem repeats in DNA. Nucleic Acids Research 31, 3672-3678.]])) ​    | Sites shorter than period + 9 are automatically discarded. ​    | Mixed combinatorial/​heuristic paradigm ​    | Start and end positions of the region to be processed; length interval, period interval, minimal exponent of the repetitions to report; resolution level, use or not of the sliding window. ​    | One file with multiple sequences in the fasta format. ​    | list of all repeats, with start and end positions of the repeat in the sequence, overall size of the repeat, period, exponent, error level, the repeat sequence itself ​    | Yes, imperfect and compound (not indels) ​    | Yes     | [[http://​bioinfo.lifl.fr/​mreps/​mreps.php|Yes]] (or [[http://​mobyle.pasteur.fr/​cgi-bin/​portal.py?#​forms::​mreps|here]]) ​    | Command-line [[http://​bioinfo.lifl.fr/​mreps//​|Download]] ​    | Linux, SunOS, Digital Unix and Windows systems. Online? ​    | ANSI C     | Should be fast. linear in the sequence length ​    | ?     | Automated statistical analysis files not generated; ​    ​| ​    | ^Mreps ​    | 2003     | ?     | [[http://​scholar.google.ca/​scholar?​cites=7808444266556719865&​as_sdt=2005&​sciodt=0,​5&​hl=fr|156]] ​    | Kolpakov et al 2003(([[http://​www.nar.oupjournals.org/​cgi/​doi/​10.1093/​nar/​gkg617|Kolpakov,​ R., Bana, G., and Kucherov, G. (2003). mreps: efficient and flexible detection of tandem repeats in DNA. Nucleic Acids Research 31, 3672-3678.]])) ​    | Sites shorter than period + 9 are automatically discarded. ​    | Mixed combinatorial/​heuristic paradigm ​    | Start and end positions of the region to be processed; length interval, period interval, minimal exponent of the repetitions to report; resolution level, use or not of the sliding window. ​    | One file with multiple sequences in the fasta format. ​    | list of all repeats, with start and end positions of the repeat in the sequence, overall size of the repeat, period, exponent, error level, the repeat sequence itself ​    | Yes, imperfect and compound (not indels) ​    | Yes     | [[http://​bioinfo.lifl.fr/​mreps/​mreps.php|Yes]] (or [[http://​mobyle.pasteur.fr/​cgi-bin/​portal.py?#​forms::​mreps|here]]) ​    | Command-line [[http://​bioinfo.lifl.fr/​mreps//​|Download]] ​    | Linux, SunOS, Digital Unix and Windows systems. Online? ​    | ANSI C     | Should be fast. linear in the sequence length ​    | ?     | Automated statistical analysis files not generated; ​    ​| ​    |
 ^STRING ​     | 2003     | 2003     | [[http://​scholar.google.ca/​scholar?​cites=7212951020758895252&​as_sdt=2005&​sciodt=0,​5&​hl=fr|32]] ​    | Parisi et al 2003(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btg268|Parisi,​ V., De Fonzo, V., and Aluffi-Pentini,​ F. (2003). STRING: finding tandem repeats in DNA sequences. Bioinformatics 19, 1733-1738.]])) ​    | No limit in repeat length ​    | Heuristic; dynamic programming procedure. ​    | ?     | One file with one sequence, no limit in length. ​    | Length of the consensus word; first and last position of the TR; number of repeated units; score; consensus word; flanking sequences; alignment between the model TR and the given sequence; number of indels; number of matches and mismatches; TR base composition percentages;​ flag indicating a likely nested expansion. ​    | Yes, imperfect and compound (and indels?​) ​    | No but may be possible ​    | A web interface is mentioned in the original publication,​ but cannot be found     | Command-line [[http://​www.caspur.it/​~castri/​STRING/​|Download]] ​    | Unix, Windows, MacOS     | C     | Slowly increases as a function of the sequence length, while it increases more quickly as a function of the number of TRs.     | ?     | Automated statistical analysis files not generated. ​    ​| ​    | ^STRING ​     | 2003     | 2003     | [[http://​scholar.google.ca/​scholar?​cites=7212951020758895252&​as_sdt=2005&​sciodt=0,​5&​hl=fr|32]] ​    | Parisi et al 2003(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btg268|Parisi,​ V., De Fonzo, V., and Aluffi-Pentini,​ F. (2003). STRING: finding tandem repeats in DNA sequences. Bioinformatics 19, 1733-1738.]])) ​    | No limit in repeat length ​    | Heuristic; dynamic programming procedure. ​    | ?     | One file with one sequence, no limit in length. ​    | Length of the consensus word; first and last position of the TR; number of repeated units; score; consensus word; flanking sequences; alignment between the model TR and the given sequence; number of indels; number of matches and mismatches; TR base composition percentages;​ flag indicating a likely nested expansion. ​    | Yes, imperfect and compound (and indels?​) ​    | No but may be possible ​    | A web interface is mentioned in the original publication,​ but cannot be found     | Command-line [[http://​www.caspur.it/​~castri/​STRING/​|Download]] ​    | Unix, Windows, MacOS     | C     | Slowly increases as a function of the sequence length, while it increases more quickly as a function of the number of TRs.     | ?     | Automated statistical analysis files not generated. ​    ​| ​    |
-^W-SSRF ​    | 2003     | ?     | [[http://​scholar.google.ca/​scholar?​cites=6592174860693683120&​as_sdt=2005&​sciodt=0,​5&​hl=fr|8]] ​    | Sreenu et al 2003(([[http://​www.ncbi.nlm.nih.gov/​pubmed/​15130803|Sreenu,​ V. B., Ranjitkumar,​ G., Swaminathan,​ S., Priya, S., Bose, B., Pavan, M. N., Thanu, G., Nagaraju, J., and Nagarajaram,​ H. A. (2003). MICAS: a fully automated web server for microsatellite extraction and analysis from prokaryote and viral genomic sequences. Appl. Bioinformatics 2, 165-168.]])) ​    | 1 to 10 bp long     | Scans a nucleotide sequence ​    | ?     | One file with one sequence. Upload limit is a 20 kb file.      | Sequence content of the motif, repeat numbers, start and end position of the tract in the sequence ​    | No, perfect only     | Yes, using Autoprimer (in MICAS) ​    | No     | The user friendly GUI (graphical user interface) is MICAS, Available upon request to the [[mailto:​[email protected]|authors]] ​    | MICAS is [[http://210.212.212.7/MIC/index.html|web-only]] ​    | Java     | ?     | ?     | ?     ​| ​    |+^W-SSRF ​    | 2003     | ?     | [[http://​scholar.google.ca/​scholar?​cites=6592174860693683120&​as_sdt=2005&​sciodt=0,​5&​hl=fr|8]] ​    | Sreenu et al 2003(([[http://​www.ncbi.nlm.nih.gov/​pubmed/​15130803|Sreenu,​ V. B., Ranjitkumar,​ G., Swaminathan,​ S., Priya, S., Bose, B., Pavan, M. N., Thanu, G., Nagaraju, J., and Nagarajaram,​ H. A. (2003). MICAS: a fully automated web server for microsatellite extraction and analysis from prokaryote and viral genomic sequences. Appl. Bioinformatics 2, 165-168.]])) ​    | 1 to 10 bp long     | Scans a nucleotide sequence ​    | ?     | One file with one sequence. Upload limit is a 20 kb file.      | Sequence content of the motif, repeat numbers, start and end position of the tract in the sequence ​    | No, perfect only     | Yes, using Autoprimer (in MICAS) ​    | No     | The user friendly GUI (graphical user interface) is MICAS, Available upon request to the [[mailto:​[email protected]|authors]] ​    | MICAS is [[http://micas.cdfd.org.in:8080/​MIC/​|web-only]] ​    | Java     | ?     | ?     | ?     ​| ​    |
 ^IRF program ​    | 2004     | 2007     | [[http://​scholar.google.ca/​scholar?​cites=7789233353140150696&​as_sdt=2005&​sciodt=0,​5&​hl=fr|80]] ​    | Warburton et al 2004(([[http://​www.genome.org/​cgi/​doi/​10.1101/​gr.2542904|Warburton,​ P. E. (2004). Inverted repeat structure of the human genome: The X-chromosome contains a preponderance of large, highly homologous inverted repeats that contain testes genes. Genome Research 14, 1861-1869.]])) ​    | ?     | ?     | ?     | ?     | ?     ​| ​     |      | ?     | Command-line [[http://​tandem.bu.edu/​news.html#​mar2007|Download]] ​    | ?     | ?     | ?     | ?     | ?     ​| ​    | ^IRF program ​    | 2004     | 2007     | [[http://​scholar.google.ca/​scholar?​cites=7789233353140150696&​as_sdt=2005&​sciodt=0,​5&​hl=fr|80]] ​    | Warburton et al 2004(([[http://​www.genome.org/​cgi/​doi/​10.1101/​gr.2542904|Warburton,​ P. E. (2004). Inverted repeat structure of the human genome: The X-chromosome contains a preponderance of large, highly homologous inverted repeats that contain testes genes. Genome Research 14, 1861-1869.]])) ​    | ?     | ?     | ?     | ?     | ?     ​| ​     |      | ?     | Command-line [[http://​tandem.bu.edu/​news.html#​mar2007|Download]] ​    | ?     | ?     | ?     | ?     | ?     ​| ​    |
 ^ExTRS ​    | 2004     | ?     | [[http://​scholar.google.ca/​scholar?​cites=10980601425678591553&​as_sdt=2005&​sciodt=0,​5&​hl=fr|33]] ​    | Krishnan and Tang 2004(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bth311|Krishnan,​ A., and Tang, F. (2004). Exhaustive whole-genome tandem repeats search. Bioinformatics 20, 2702-2710.]])) ​    | ?     | Exhaustive ​    | ?     | One file with one sequence, no limit in length. ​    | Redundancy in the output is reduced ​    | Yes, substitutions (not indels) ​    ​| ​     | ?     | ? Available upon request to the [[mailto:​[email protected]|authors]] ​    | ?     | Source code available only on request ​    | Near-proportional to the number of TR found     | ?     | ?     ​| ​    | ^ExTRS ​    | 2004     | ?     | [[http://​scholar.google.ca/​scholar?​cites=10980601425678591553&​as_sdt=2005&​sciodt=0,​5&​hl=fr|33]] ​    | Krishnan and Tang 2004(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bth311|Krishnan,​ A., and Tang, F. (2004). Exhaustive whole-genome tandem repeats search. Bioinformatics 20, 2702-2710.]])) ​    | ?     | Exhaustive ​    | ?     | One file with one sequence, no limit in length. ​    | Redundancy in the output is reduced ​    | Yes, substitutions (not indels) ​    ​| ​     | ?     | ? Available upon request to the [[mailto:​[email protected]|authors]] ​    | ?     | Source code available only on request ​    | Near-proportional to the number of TR found     | ?     | ?     ​| ​    |
Line 30: Line 30:
 ^TRA     | 2004     | 2004     | [[http://​scholar.google.ca/​scholar?​cites=13056073161710952438&​as_sdt=5&​sciodt=0&​hl=fr|21]] ​    | Bilgen et al 2004(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bth410|Bilgen,​ M., Karaca, M., Onus, A. N., and Ince, A. G. (2004). A software program combining sequence motif searches with keywords for finding repeats containing DNA sequences. Bioinformatics 20, 3379-3386.]])) ​    | ?     | Heuristic ​    | ?     | Multiple files with multiple sequences each (Max 1 Mb sequence) from ESTs.     | ?     | Yes, searches for exact–inexact TRs and exact–inexact compound repeats ​    | No?     | ?     | ? [[ftp://​ftp.akdeniz.edu.tr/​Araclar/​TRA/​|Download]] ​    | Windows ​    | C++ (with Microsoft Visual C++)     | ?     | Searches among the organisms, organs, tissue types and development stages ​    | ?     ​| ​    | ^TRA     | 2004     | 2004     | [[http://​scholar.google.ca/​scholar?​cites=13056073161710952438&​as_sdt=5&​sciodt=0&​hl=fr|21]] ​    | Bilgen et al 2004(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bth410|Bilgen,​ M., Karaca, M., Onus, A. N., and Ince, A. G. (2004). A software program combining sequence motif searches with keywords for finding repeats containing DNA sequences. Bioinformatics 20, 3379-3386.]])) ​    | ?     | Heuristic ​    | ?     | Multiple files with multiple sequences each (Max 1 Mb sequence) from ESTs.     | ?     | Yes, searches for exact–inexact TRs and exact–inexact compound repeats ​    | No?     | ?     | ? [[ftp://​ftp.akdeniz.edu.tr/​Araclar/​TRA/​|Download]] ​    | Windows ​    | C++ (with Microsoft Visual C++)     | ?     | Searches among the organisms, organs, tissue types and development stages ​    | ?     ​| ​    |
 ^MsatFinder ​    | 2005     | 2007     | 48     | Thurston and Field 2005(([[http://​www.genomics.ceh.ac.uk/​msatfinder/​|Thurston,​ M. I., and Field, D. (2006). Msatfinder, (Oxford, UK: Centre for Ecology and Hydrology. Computer program)]])) ​    | One to 6 bp long     | ?     | 1) length of repeat 2) number repeat unit in the site 3) search engine (regex, multipass or iterative search) ​    | Limit of 10 Mb of sequence in the online access. Accepts GenBank, EMBL, Swissprot, FASTA, ASCII. ​    | Repeats, GFF, Counts, Msat_tabs, Flank_tabs, Fasta, MINE, Primers ​    | No, but detects compound perfect repeats ​    ​| ​     | [[http://​www.genomics.ceh.ac.uk/​cgi-bin/​msatfinder/​msatfinder.cgi|Yes]] ​    | Command-line [[http://​www.genomics.ceh.ac.uk/​msatfinder/#​download|Download]] ​    | Unix (may work on Mac OSX)     | perl script ​    | ?     | Nucleic acid or amino acid sequence ​    | ?     ​| ​    | ^MsatFinder ​    | 2005     | 2007     | 48     | Thurston and Field 2005(([[http://​www.genomics.ceh.ac.uk/​msatfinder/​|Thurston,​ M. I., and Field, D. (2006). Msatfinder, (Oxford, UK: Centre for Ecology and Hydrology. Computer program)]])) ​    | One to 6 bp long     | ?     | 1) length of repeat 2) number repeat unit in the site 3) search engine (regex, multipass or iterative search) ​    | Limit of 10 Mb of sequence in the online access. Accepts GenBank, EMBL, Swissprot, FASTA, ASCII. ​    | Repeats, GFF, Counts, Msat_tabs, Flank_tabs, Fasta, MINE, Primers ​    | No, but detects compound perfect repeats ​    ​| ​     | [[http://​www.genomics.ceh.ac.uk/​cgi-bin/​msatfinder/​msatfinder.cgi|Yes]] ​    | Command-line [[http://​www.genomics.ceh.ac.uk/​msatfinder/#​download|Download]] ​    | Unix (may work on Mac OSX)     | perl script ​    | ?     | Nucleic acid or amino acid sequence ​    | ?     ​| ​    |
-^Fireusat ​    | 2006     | 2011     | [[http://​scholar.google.ca/​scholar?​cites=17060705618169575807&​as_sdt=2005&​sciodt=0,​5&​hl=fr|5]] and [[http://​scholar.google.ca/​scholar?​cites=66353776030260478&​as_sdt=5&​sciodt=0&​hl=fr|this]] ​    | de Ridder at al 2006(([[http://​portal.acm.org/​citation.cfm?​doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). ​FireμSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing countries ​ - SAICSIT ​ ’06 (Somerset West, South Africa), pp. 247-256.]])) and [[http://upetd.up.ac.za/thesis/available/etd-08172010-202532/unrestricted/​dissertation.pdf|de Ridder ​2006]]     | 1 to 5 bp. Length set by the user, usually > 2 bp long     | Uses Counting Finite Automata (which are regular language acceptors) ​    | Max Motif Error (per motif); Max adjacent ATR elements; Motif Range Options; Min required TR elements; Max substring error (a threshold); Mismatch penalty (m_p); Delete penalty (d_p); Insert penalty (i_p). ​    | One fasta file with one sequence ​    | File in .csv format. ​    | Yes, substitutions and indels, but not compound loci.     | No     | No     | GUI and command-line,​ [[http://​www.dna-algo.co.za/​|Download]] ​    | Windows; Linux in progress ​    | C++ and MatLab ​    | Run time increases linearly with the sequence length; does not increase with longer motif lengths. ​    | Designed for microsatellites,​ but can detect any type of TR     | Fast, simple and flexible. ​    ​| ​    |+^FireµSat ​    | 2006     | 2011     | [[http://​scholar.google.ca/​scholar?​cites=17060705618169575807&​as_sdt=2005&​sciodt=0,​5&​hl=fr|5]] and [[http://​scholar.google.ca/​scholar?​cites=66353776030260478&​as_sdt=5&​sciodt=0&​hl=fr|this]] ​    | de Ridder at al 2006(([[http://​portal.acm.org/​citation.cfm?​doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). ​FireµSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing countries ​ - SAICSIT ​ ’06 (Somerset West, South Africa), pp. 247-256.]])) and de Ridder at al 2013(([[http://www.sciencedirect.com/science/article/pii/S1570866712001657|De Ridder, C., D.G. Kourie, B.W. Watson, T.R. Fourie, and P.V. Reyneke (2013). Fine-tuning the search for microsatellites. Journal of Discrete Algorithms 20: 21–37.]]))     | 1 to 5 bp. Length set by the user. The next update should allow for detection of 6 to 100 bp repeats. ​    | Uses Counting Finite Automata (which are regular language acceptors) ​    | Max Motif Error (per motif); Max adjacent ATR elements; Motif Range Options; Min required TR elements; Max substring error (a threshold); Mismatch penalty (m_p); Delete penalty (d_p); Insert penalty (i_p). ​    | One fasta file with one sequence ​    | File in .csv format. ​    | Yes, substitutions and indels, but not compound loci.     | No     | No     | GUI and command-line,​ [[http://​www.dna-algo.co.za/​|Download]] ​    | Windows; Linux in progress ​    | C++ and MatLab ​    | Run time increases linearly with the sequence length; does not increase with longer motif lengths. ​    | Designed for microsatellites,​ but can detect any type of TR     | Fast, simple and flexible. ​    ​| ​    |
 ^Phobos ​    | 2006     | 2010     | ?     | Mayer 2010(([[http://​www.ruhr-uni-bochum.de/​spezzoo/​cm/​cm_phobos.htm |Mayer, C. (2010). Phobos: Highly accurate search for perfect and imperfect tandem repeats in complete genomes by Christoph Mayer, (Bochum, Germany: Ruhr-Universität Bochum,​Faculty of Biological Sciences and Biotechnology). Computer program.]])) ​     | Perfect and imperfect TR, with a pattern size of 1 - 10 000 bp     | Exhaustive, uses alignment scores ​    | Mismatch score, indel score, minimum score, minimum length, minimum perfection, and others. ​    | One file in fasta format, with multiple sequence. No limit in sequence length. ​    | Text file, different formats, including gff and fasta     | Yes. Substitutions and indels. ​     | Not in itself, but yes as implemented in STAMP or Geneious. ​     | No     | User friendly GUI and easily scriptable Command-line program [[http://​www.ruhr-uni-bochum.de/​ecoevo/​cm/​cm_phobos.htm|Download]] ​    | MacOSX, Linux, Windows. ​    | C++     | Execution time increases with pattern size range. Very fast in the size range 1-10 bp, slow for patterns in the size range above 10-20 bp.     | Can be incorporated into pipelines. Implemented in STAMP and Geneious. ​    | Free only for academic users. ​    ​| ​    | ^Phobos ​    | 2006     | 2010     | ?     | Mayer 2010(([[http://​www.ruhr-uni-bochum.de/​spezzoo/​cm/​cm_phobos.htm |Mayer, C. (2010). Phobos: Highly accurate search for perfect and imperfect tandem repeats in complete genomes by Christoph Mayer, (Bochum, Germany: Ruhr-Universität Bochum,​Faculty of Biological Sciences and Biotechnology). Computer program.]])) ​     | Perfect and imperfect TR, with a pattern size of 1 - 10 000 bp     | Exhaustive, uses alignment scores ​    | Mismatch score, indel score, minimum score, minimum length, minimum perfection, and others. ​    | One file in fasta format, with multiple sequence. No limit in sequence length. ​    | Text file, different formats, including gff and fasta     | Yes. Substitutions and indels. ​     | Not in itself, but yes as implemented in STAMP or Geneious. ​     | No     | User friendly GUI and easily scriptable Command-line program [[http://​www.ruhr-uni-bochum.de/​ecoevo/​cm/​cm_phobos.htm|Download]] ​    | MacOSX, Linux, Windows. ​    | C++     | Execution time increases with pattern size range. Very fast in the size range 1-10 bp, slow for patterns in the size range above 10-20 bp.     | Can be incorporated into pipelines. Implemented in STAMP and Geneious. ​    | Free only for academic users. ​    ​| ​    |
 ^SSRscanner ​    | 2006     | ?     ​| ​ [[http://​scholar.google.ca/​scholar?​cites=15246530144073441813&​as_sdt=2005&​sciodt=0,​5&​hl=fr|3]] ​    | Anwar and Khan 2006(([[http://​www.ncbi.nlm.nih.gov/​pmc/​articles/​PMC1891659/​|Anwar,​ T., and Khan, A. (2006). SSRscanner: a program for reporting distribution and exact location of simple sequence repeats. Bioinformation 1, 89-91.]])) ​    | Only searches for predefined motifs ​    | Exhaustive, uses dictionary approach. ​    | File containing motifs of different repeat types; number of times for the motifs to be repeated ​    | One file with one sequence ​    | Motifposition.txt (gives the frequency of each repeat provided in the motif file) and (2) Motifresult.exe (gives the specific location of each repeat) ​    | No, perfect only     | No     | No     | Command-line Availability unknown, contact the [[mailto:​[email protected]|author]] ​    | Platform independent ​    | perl script ​    | ?     | ?     | ?     ​| ​    | ^SSRscanner ​    | 2006     | ?     ​| ​ [[http://​scholar.google.ca/​scholar?​cites=15246530144073441813&​as_sdt=2005&​sciodt=0,​5&​hl=fr|3]] ​    | Anwar and Khan 2006(([[http://​www.ncbi.nlm.nih.gov/​pmc/​articles/​PMC1891659/​|Anwar,​ T., and Khan, A. (2006). SSRscanner: a program for reporting distribution and exact location of simple sequence repeats. Bioinformation 1, 89-91.]])) ​    | Only searches for predefined motifs ​    | Exhaustive, uses dictionary approach. ​    | File containing motifs of different repeat types; number of times for the motifs to be repeated ​    | One file with one sequence ​    | Motifposition.txt (gives the frequency of each repeat provided in the motif file) and (2) Motifresult.exe (gives the specific location of each repeat) ​    | No, perfect only     | No     | No     | Command-line Availability unknown, contact the [[mailto:​[email protected]|author]] ​    | Platform independent ​    | perl script ​    | ?     | ?     | ?     ​| ​    |
Line 38: Line 38:
 ^SciRoKo ​    | 2007     | 2008     | [[http://​scholar.google.ca/​scholar?​cites=16996584693111364287&​as_sdt=2005&​sciodt=0,​5&​hl=fr|44]] ​    | Kofler et al 2007(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btm157|Kofler,​ R., Schlotterer,​ C., and Lelley, T. (2007). SciRoKo: a new tool for whole genome microsatellite search and investigation. Bioinformatics 23, 1683-1685.]])) ​    | One to 6 bp long     | ?     | Hits (identity with a virtual perfect microsatellite),​ number of mismatches (mm), mismatch penalty (mmP) and the length of the SSR motif (mL).     | One file with multiple sequences in the fasta format. ​    | ?     | Yes, imperfect and compound (indels?​) ​    ​| ​     | No     | User friendly GUI (graphical user interface) standalone. [[http://​www.kofler.or.at/​bioinformatics/​SciRoKo/​index.html|Download]] ​    | Windows. Should be platform independent,​ but Mac users have not been able install ​    | C#     | Fast     | ?     | Depends on .NET framework ​    ​| ​    | ^SciRoKo ​    | 2007     | 2008     | [[http://​scholar.google.ca/​scholar?​cites=16996584693111364287&​as_sdt=2005&​sciodt=0,​5&​hl=fr|44]] ​    | Kofler et al 2007(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btm157|Kofler,​ R., Schlotterer,​ C., and Lelley, T. (2007). SciRoKo: a new tool for whole genome microsatellite search and investigation. Bioinformatics 23, 1683-1685.]])) ​    | One to 6 bp long     | ?     | Hits (identity with a virtual perfect microsatellite),​ number of mismatches (mm), mismatch penalty (mmP) and the length of the SSR motif (mL).     | One file with multiple sequences in the fasta format. ​    | ?     | Yes, imperfect and compound (indels?​) ​    ​| ​     | No     | User friendly GUI (graphical user interface) standalone. [[http://​www.kofler.or.at/​bioinformatics/​SciRoKo/​index.html|Download]] ​    | Windows. Should be platform independent,​ but Mac users have not been able install ​    | C#     | Fast     | ?     | Depends on .NET framework ​    ​| ​    |
 ^Msatcommander ​    | 2008     | 2011     | [[http://​scholar.google.ca/​scholar?​cites=3185583603820687905&​as_sdt=2005&​sciodt=0,​5&​hl=fr|102]] ​    | Faircloth 2008(([[http://​doi.wiley.com/​10.1111/​j.1471-8286.2007.01884.x|Faircloth,​ B. C. (2008). msatcommander:​ detection of microsatellite repeat arrays and automated, locus-specific primer design. Molecular Ecology Resources 8, 92-94.]])) ​    | ?     | Uses regular expressions ​    | ?     | One file with multiple sequences in the fasta format. ​    | Either a summary file (array detection only) or a directory at a user-selectable location ​    | No, but accepts N     | Yes, with Primer3, includes 5'​-tailing ​    | No     | User friendly GUI (graphical user interface) standalone. [[http://​code.google.com/​p/​msatcommander/​|Download]] ​    | MacOS X, Windows, Unix.     | Python ​    | ?     | Rapid and automated microsatellite array detection, locus-specific primer design, and 5'​-tailing of designed primers ​    | ?     ​| ​    | ^Msatcommander ​    | 2008     | 2011     | [[http://​scholar.google.ca/​scholar?​cites=3185583603820687905&​as_sdt=2005&​sciodt=0,​5&​hl=fr|102]] ​    | Faircloth 2008(([[http://​doi.wiley.com/​10.1111/​j.1471-8286.2007.01884.x|Faircloth,​ B. C. (2008). msatcommander:​ detection of microsatellite repeat arrays and automated, locus-specific primer design. Molecular Ecology Resources 8, 92-94.]])) ​    | ?     | Uses regular expressions ​    | ?     | One file with multiple sequences in the fasta format. ​    | Either a summary file (array detection only) or a directory at a user-selectable location ​    | No, but accepts N     | Yes, with Primer3, includes 5'​-tailing ​    | No     | User friendly GUI (graphical user interface) standalone. [[http://​code.google.com/​p/​msatcommander/​|Download]] ​    | MacOS X, Windows, Unix.     | Python ​    | ?     | Rapid and automated microsatellite array detection, locus-specific primer design, and 5'​-tailing of designed primers ​    | ?     ​| ​    |
-^ReRep ​    | 2008     | 2008     | [[http://​scholar.google.ca/​scholar?​cites=255635357253003447&​as_sdt=2005&​sciodt=0,​5&​hl=fr|5]] ​    | Otto et al 2008(([[http://​www.biomedcentral.com/​1471-2105/​9/​366|Otto,​ T. D., Gomes, L. H. F., Alves-Ferreira,​ M., de Miranda, A. B., and Degrave, W. M. (2008). ReRep: Computational detection of repetitive sequences in genome survey sequences (GSS). BMC Bioinformatics 9, 366.]])) ​    | ?     | Uses self-similarity searches ​    | ?     | Genome survey sequences(GSS) files, including 454-reads ​    | ?     | Yes, substitutions and indels ​    | No     | No     | Command-line [[http://bioinfo.pdtis.fiocruz.br/​ReRep/​|Download]] ​    | Linux     | Perl     | ?     | Can detect de novo repeats in Genome Sequence Survey sequence data     | ?    |+^ReRep ​    | 2008     | 2008     | [[http://​scholar.google.ca/​scholar?​cites=255635357253003447&​as_sdt=2005&​sciodt=0,​5&​hl=fr|5]] ​    | Otto et al 2008(([[http://​www.biomedcentral.com/​1471-2105/​9/​366|Otto,​ T. D., Gomes, L. H. F., Alves-Ferreira,​ M., de Miranda, A. B., and Degrave, W. M. (2008). ReRep: Computational detection of repetitive sequences in genome survey sequences (GSS). BMC Bioinformatics 9, 366.]])) ​    | ?     | Uses self-similarity searches ​    | ?     | Genome survey sequences(GSS) files, including 454-reads ​    | ?     | Yes, substitutions and indels ​    | No     | No     | Command-line [[http://www.dbbm.fiocruz.br/​labwim/bioinfoteam/​index.pl?​action=services|Download]] ​    | Linux     | Perl     | ?     | Can detect de novo repeats in Genome Sequence Survey sequence data     | ?    |
 ^T-REKS ​    | 2009     | ?     | [[http://​scholar.google.ca/​scholar?​cites=208721574341442296&​as_sdt=2005&​sciodt=0,​5&​hl=fr|5]] ​    | Jorda and Kajava 2009(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btp482|Jorda,​ J., and Kajava, A. V. (2009). T-REKS: identification of Tandem REpeats in sequences with a K-meanS based algorithm. Bioinformatics 25, 2632-2638.]])) ​    | No limits in repeat length ​    | Short string extension and K-means algorithm. ​    | delta-l Allowed % of length variability,​ P*sim—similarity threshold and an option to allow or not  the detection of overlaping TR.     | Sequences in FASTA format. ​    | Output with start, end, length of TR and multiple alignment of the repeats. ​    | Yes, substitutions and indels. ​    | No     | [[http://​bioinfo.montp.cnrs.fr/?​r=t-reks|Yes]] ​    | User friendly GUI (graphical user interface) standalone. [[http://​bioinfo.montp.cnrs.fr/?​r=t-reks|Download]] ​    | Platform independent ​    | Java     | Fast. Execution time is linear (directly proportional) to the sequence length. ​    | Can be applied to nucleic acid, amino acid or any text sequence. ​    | ?    | ^T-REKS ​    | 2009     | ?     | [[http://​scholar.google.ca/​scholar?​cites=208721574341442296&​as_sdt=2005&​sciodt=0,​5&​hl=fr|5]] ​    | Jorda and Kajava 2009(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btp482|Jorda,​ J., and Kajava, A. V. (2009). T-REKS: identification of Tandem REpeats in sequences with a K-meanS based algorithm. Bioinformatics 25, 2632-2638.]])) ​    | No limits in repeat length ​    | Short string extension and K-means algorithm. ​    | delta-l Allowed % of length variability,​ P*sim—similarity threshold and an option to allow or not  the detection of overlaping TR.     | Sequences in FASTA format. ​    | Output with start, end, length of TR and multiple alignment of the repeats. ​    | Yes, substitutions and indels. ​    | No     | [[http://​bioinfo.montp.cnrs.fr/?​r=t-reks|Yes]] ​    | User friendly GUI (graphical user interface) standalone. [[http://​bioinfo.montp.cnrs.fr/?​r=t-reks|Download]] ​    | Platform independent ​    | Java     | Fast. Execution time is linear (directly proportional) to the sequence length. ​    | Can be applied to nucleic acid, amino acid or any text sequence. ​    | ?    |
 ^BwTRS ​    | 2010     | 2009     | [[http://​scholar.google.ca/​scholar?​cites=7603849771779688492&​as_sdt=2005&​sciodt=0,​5&​hl=fr|2]] ​    | Pokrzywa and Polanski 2010(([[http://​linkinghub.elsevier.com/​retrieve/​pii/​S0888754310001758|Pokrzywa,​ R., and Polanski, A. (2010). BWtrs: A tool for searching for tandem repeats in DNA sequences based on the Burrows–Wheeler transform. Genomics 96, 316-321.]])) ​    | ?     | Exhaustive; uses efficient data compression algorithm ​    | "​Minimum motif size", "​Maximum motif size", "​Minimum repeat size" and "​Minimum repeat ratio"​. ​    | One file with multiple sequences in the fasta format; or GenBank id. Accepts nucleotides and amino acids sequences. ​    | List of all TR with: Start and End (position of the TR in the sequence); Motif length; Ratio (between the motif length and the consensus repeat length); the motif itself. HTML or text.     | No     | No     | [[http://​bwtools.polsl.pl/​BWtrs/​input.jsp|Yes]],​ runs on a standard Tomcat servlet container without any local database ​    | ? Availability unknown, contact the [[mailto: [email protected]|authors]]. ​    | ?     | Java     | Depends on sequence length ​    | Nucleic acid or amino acid sequence ​    | ?    | ^BwTRS ​    | 2010     | 2009     | [[http://​scholar.google.ca/​scholar?​cites=7603849771779688492&​as_sdt=2005&​sciodt=0,​5&​hl=fr|2]] ​    | Pokrzywa and Polanski 2010(([[http://​linkinghub.elsevier.com/​retrieve/​pii/​S0888754310001758|Pokrzywa,​ R., and Polanski, A. (2010). BWtrs: A tool for searching for tandem repeats in DNA sequences based on the Burrows–Wheeler transform. Genomics 96, 316-321.]])) ​    | ?     | Exhaustive; uses efficient data compression algorithm ​    | "​Minimum motif size", "​Maximum motif size", "​Minimum repeat size" and "​Minimum repeat ratio"​. ​    | One file with multiple sequences in the fasta format; or GenBank id. Accepts nucleotides and amino acids sequences. ​    | List of all TR with: Start and End (position of the TR in the sequence); Motif length; Ratio (between the motif length and the consensus repeat length); the motif itself. HTML or text.     | No     | No     | [[http://​bwtools.polsl.pl/​BWtrs/​input.jsp|Yes]],​ runs on a standard Tomcat servlet container without any local database ​    | ? Availability unknown, contact the [[mailto: [email protected]|authors]]. ​    | ?     | Java     | Depends on sequence length ​    | Nucleic acid or amino acid sequence ​    | ?    |
Line 46: Line 46:
 ^repeatfinder ​    | 2001     | ?     | ?     | Volfovsky et al 2001(([[http://​genomebiology.com/​2001/​2/​8/​research/​0027|Volfovsky,​ N., Haas, B., and Salzberg, S. (2001). A clustering method for repeat analysis in DNA sequences. Genome Biology 2, research0027.1 - research0027.11.]])) ​    | ?     | Uses REPuter ​    | ?     | ?     | ?     | ?     | ?     | No     | Command-line [[http://​www.cbcb.umd.edu/​software/​RepeatFinder/​|Download]] ​    | Linux RedHat 6.x+, Sun Solaris, and Alpha OSF1     | ? Open Source ​    | ?     | ?     | ?    | ^repeatfinder ​    | 2001     | ?     | ?     | Volfovsky et al 2001(([[http://​genomebiology.com/​2001/​2/​8/​research/​0027|Volfovsky,​ N., Haas, B., and Salzberg, S. (2001). A clustering method for repeat analysis in DNA sequences. Genome Biology 2, research0027.1 - research0027.11.]])) ​    | ?     | Uses REPuter ​    | ?     | ?     | ?     | ?     | ?     | No     | Command-line [[http://​www.cbcb.umd.edu/​software/​RepeatFinder/​|Download]] ​    | Linux RedHat 6.x+, Sun Solaris, and Alpha OSF1     | ? Open Source ​    | ?     | ?     | ?    |
 ^MsatMiner ​    | 2005     | ?     ​| ​     | Thurston and Field 2005(([[http://​www.genomics.ceh.ac.uk/​msatminer/​|Thurston,​ M. I. (2005). Msatminer - scripts for processing msatfinder output, (Oxford, UK: Centre for Ecology and Hydrology). Computer program.]])) ​    | ?     | Uses msatfinder for motif discovery ​    ​| ​     | Different format of sequence file     | More statistical analysis are possible. ​    | Yes, imperfect and compound (indels?​) ​    | Possible when using additional scripts ​    | ?     | Command-line ​    | Unix and MacOS     | Collection of perl scripts ​    | ?     | ?     | Running scripts is possibly complicated ​   | ^MsatMiner ​    | 2005     | ?     ​| ​     | Thurston and Field 2005(([[http://​www.genomics.ceh.ac.uk/​msatminer/​|Thurston,​ M. I. (2005). Msatminer - scripts for processing msatfinder output, (Oxford, UK: Centre for Ecology and Hydrology). Computer program.]])) ​    | ?     | Uses msatfinder for motif discovery ​    ​| ​     | Different format of sequence file     | More statistical analysis are possible. ​    | Yes, imperfect and compound (indels?​) ​    | Possible when using additional scripts ​    | ?     | Command-line ​    | Unix and MacOS     | Collection of perl scripts ​    | ?     | ?     | Running scripts is possibly complicated ​   |
-^E-TRA ​    | 2005     | 2004     | [[http://​scholar.google.ca/​scholar?​cites=6055835177193667419&​as_sdt=5&​sciodt=0&​hl=fr|12]] ​    | Karaca et al 2005(([[Available at: http://​www.springerlink.com/​index/​10.1007/​BF02715889|Karaca,​ M., Bilgen, M., Onus, A. N., Ince, A. G., and Elmasulu, S. Y. (2005). Exact tandem repeats analyzer (E-TRA): A new program for DNA sequence mining. J Genet 84, 49-54.]])) ​    | 1 to 1000 bp repeat ​    | Uses TRA     | ?     | Multiple files with multiple sequences each (maximum of 1 Mb long)     ​| ​     | Yes, compound and imperfect ​    | Yes     | No     | user friendly GUI (graphical user interface) [[ftp://​ftp.akdeniz.edu.tr/​Araclar/​TRA/​|Download]] ​    | Windows ​    | C++ (with Microsoft Visual C++)     ​| ​     | Searches among the organisms, organs, tissue types and development stages ​    | Only 1 Mb of input sequence length ​   | +^E-TRA ​    | 2005     | 2004     | [[http://​scholar.google.ca/​scholar?​cites=6055835177193667419&​as_sdt=5&​sciodt=0&​hl=fr|12]] ​    | Karaca et al 2005(([[http://​www.springerlink.com/​index/​10.1007/​BF02715889|Karaca,​ M., Bilgen, M., Onus, A. N., Ince, A. G., and Elmasulu, S. Y. (2005). Exact tandem repeats analyzer (E-TRA): A new program for DNA sequence mining. J Genet 84, 49-54.]])) ​    | 1 to 1000 bp repeat ​    | Uses TRA     | ?     | Multiple files with multiple sequences each (maximum of 1 Mb long)     ​| ​     | Yes, compound and imperfect ​    | Yes     | No     | user friendly GUI (graphical user interface) [[ftp://​ftp.akdeniz.edu.tr/​Araclar/​TRA/​|Download]] ​    | Windows ​    | C++ (with Microsoft Visual C++)     ​| ​     | Searches among the organisms, organs, tissue types and development stages ​    | Only 1 Mb of input sequence length ​   | 
-^SSRprimerII ​    | 2006     | 2009     | 29     | Robinson et al. 2004(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bth104|Robinson,​ A. J., Love, C. G., Batley, J., Barker, G., and Edwards, D. (2004). Simple sequence repeat marker loci discovery using SSR primer. Bioinformatics 20, 1475-1476.]])) and Jewell et al. 2006(([[http://​www.nar.oxfordjournals.org/​cgi/​doi/​10.1093/​nar/​gkl083|Jewell,​ E. et al. (2006). SSRPrimer and SSR Taxonomy Tree: Biome SSR discovery. Nucleic Acids Research 34, W656-W659.]])) ​    | 2 to >6 bp long repeats ​    | Uses Sputnik ​    ​| ​     | Limit of 4000 bp sequence ​    ​| ​     |      | Yes uses Primer3 ​    | [[http://​flora.acpfg.com.au/​ssrprimer2/​cgi-bin/​index|Yes]] ​    | ? Contact the [[ttp://​www.appliedbioinformatics.com.au/​|group]] ​    | ?     | Perl scripts ​    | ?     | ?     | ?    | +^SSRprimerII ​    | 2006     | 2009     | 29     | Robinson et al. 2004(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bth104|Robinson,​ A. J., Love, C. G., Batley, J., Barker, G., and Edwards, D. (2004). Simple sequence repeat marker loci discovery using SSR primer. Bioinformatics 20, 1475-1476.]])) and Jewell et al. 2006(([[http://​www.nar.oxfordjournals.org/​cgi/​doi/​10.1093/​nar/​gkl083|Jewell,​ E. et al. (2006). SSRPrimer and SSR Taxonomy Tree: Biome SSR discovery. Nucleic Acids Research 34, W656-W659.]])) ​    | 2 to >6 bp long repeats ​    | Uses Sputnik ​    ​| ​     | Limit of 4000 bp sequence ​    ​| ​     |      | Yes uses Primer3 ​    | [[http://​flora.acpfg.com.au/​ssrprimer2/​cgi-bin/​index|Yes]] ​    | ? Contact the [[http://​www.appliedbioinformatics.com.au/​|group]] ​    | ?     | Perl scripts ​    | ?     | ?     | ?    | 
-^TRAP     | 2006     | 2005     | 10     | Sobreira et al 2006(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bti809|Sobreira,​ T. J. P., Durham, A. M., and Gruber, A. (2005). TRAP: automated classification,​ quantification and annotation of tandemly repeated sequences. Bioinformatics 22, 361-362.]])) ​    | 1 to 2000 bp repeat ​  | Uses TRF     | All TRF parameters: ​minimum ​and maximum ​number of motifs, minimum ​and maximum ​motif size (period), ​minimum ​size of flanking regions and minimum ​match % between adjacent motifs. ​    | One file with multiple sequences in fasta format. ​    ​| ​Output files in many format ​(csv, HTML, flat files, and GFF)     ​| ​  Yes, imperfect and compound ​  ​|  ​no    | no     | Command-line [[http://​www.coccidia.icb.usp.br/​trap/​trapRegister/​|Download]] ​    | Unix     | Perl scripts ​    ​| ​ Near-proportional to the number of TR founded ​by TRF   | Selection, classification,​ quantification and automated annotation of TR sequences ​    ​| ​    |+^TRAP     | 2006     | 2005     | 10     | Sobreira et al 2006(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​bti809|Sobreira,​ T. J. P., Durham, A. M., and Gruber, A. (2005). TRAP: automated classification,​ quantification and annotation of tandemly repeated sequences. Bioinformatics 22, 361-362.]])) ​    | 1 to 2000 bp repeat ​  | Uses TRF     | All TRF parameters, as well asmin and max number of motifs; min and max motif size (period), ​min size of flanking regions and min match % between adjacent motifs. ​    | One file with multiple sequences in fasta format. ​    ​| ​Many formats ​(csv, HTML, flat files, and GFF)     ​| ​  Yes, imperfect and compound ​  ​|  ​No, but generates a fasta file with TR regions masked with Ns    | no     | Command-line [[http://​www.coccidia.icb.usp.br/​trap/​trapRegister/​|Download]] ​    | Unix; MacOSX ​    | Perl scripts ​    ​| ​ Near-proportional to the number of TRs found by TRF   | Selection, classification,​ quantification and automated annotation of TR sequences ​    ​| ​Not compatible with MS Windows version of TRF     |
 ^cid     | 2008     | ?     | 11     | Freita et al 2008(([[http://​doi.wiley.com/​10.1111/​j.1471-8286.2007.01950.x|Freitas,​ P. D., Martins, D. S., and Galetti, P. M. (2008). cid: a rapid and efficient bioinformatic tool for the detection of SSRs from genomic libraries. Molecular Ecology Resources 8, 107-108.]])) ​    | Same as MISA     | Uses MISA for tandem repeat detection, and other external programs for other steps     ​| ​     | Set of chromatograms or multiFASTA file     | List of useful primers ​    ​| ​     | Yes, using Primer3 ​    | ?     | Web environment,​ Availability unknown, contact the [[mailto: [email protected]|authors]] ​    | ?     | perl and php to connect the different tools     | ?     | Can mask vectors and adaptors regions of cloned sequences ​    | ?    | ^cid     | 2008     | ?     | 11     | Freita et al 2008(([[http://​doi.wiley.com/​10.1111/​j.1471-8286.2007.01950.x|Freitas,​ P. D., Martins, D. S., and Galetti, P. M. (2008). cid: a rapid and efficient bioinformatic tool for the detection of SSRs from genomic libraries. Molecular Ecology Resources 8, 107-108.]])) ​    | Same as MISA     | Uses MISA for tandem repeat detection, and other external programs for other steps     ​| ​     | Set of chromatograms or multiFASTA file     | List of useful primers ​    ​| ​     | Yes, using Primer3 ​    | ?     | Web environment,​ Availability unknown, contact the [[mailto: [email protected]|authors]] ​    | ?     | perl and php to connect the different tools     | ?     | Can mask vectors and adaptors regions of cloned sequences ​    | ?    |
 ^Etandem (in EMBOSS) ​    | 2008     | ?     | [[http://​scholar.google.ca/​scholar?​cites=2718801218157611554&​as_sdt=2005&​sciodt=0,​5&​hl=fr|2119]] ​    | Rice et al 2010(([[http://​linkinghub.elsevier.com/​retrieve/​pii/​S0168952500020242|Rice,​ P., Longden, I., and Bleasby, A. (2000). EMBOSS: The European Molecular Biology Open Software Suite. Trends in Genetics 16, 276-277.]])) ​    | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?    | ^Etandem (in EMBOSS) ​    | 2008     | ?     | [[http://​scholar.google.ca/​scholar?​cites=2718801218157611554&​as_sdt=2005&​sciodt=0,​5&​hl=fr|2119]] ​    | Rice et al 2010(([[http://​linkinghub.elsevier.com/​retrieve/​pii/​S0168952500020242|Rice,​ P., Longden, I., and Bleasby, A. (2000). EMBOSS: The European Molecular Biology Open Software Suite. Trends in Genetics 16, 276-277.]])) ​    | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?     | ?    |
Line 125: Line 125:
 Thurston and Field 2005(([[http://​www.genomics.ceh.ac.uk/​msatfinder/​|Thurston,​ M. I., and Field, D. (2006). Msatfinder, (Oxford, UK: Centre for Ecology and Hydrology. Computer program)]])) No published description of the algorithm could be found. According to the online access, regular expression (regex), multipass or iterative search engines can be used. Thurston and Field 2005(([[http://​www.genomics.ceh.ac.uk/​msatfinder/​|Thurston,​ M. I., and Field, D. (2006). Msatfinder, (Oxford, UK: Centre for Ecology and Hydrology. Computer program)]])) No published description of the algorithm could be found. According to the online access, regular expression (regex), multipass or iterative search engines can be used.
  
-==Fireusat==  +==FireµSat==  
-de Riddera ​at al 2006(([[http://​portal.acm.org/​citation.cfm?​doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). ​FireμSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing couuntries ​ - SAICSIT ​ ’06 (Somerset West, South Africa), pp. 247-256.]])) Is a combination of straightforward FA technology combined with a flavour of Moore machine technology. It uses regular ​expressions+de Ridder ​at al 2006(([[http://​portal.acm.org/​citation.cfm?​doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). ​FireµSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing couuntries ​ - SAICSIT ​ ’06 (Somerset West, South Africa), pp. 247-256.]])) ​and de Ridder at al 2013(([[http://​www.sciencedirect.com/​science/​article/​pii/​S1570866712001657|De Ridder, C., D.G. Kourie, B.W. Watson, T.R. Fourie, and P.V. Reyneke (2013). Fine-tuning the search for microsatellites. Journal of Discrete Algorithms 20: 21–37.]])), ​Is a combination of straightforward FA technology combined with a flavour of Moore machine technology. It uses Counting Finite Automata, which are regular ​language acceptorsThe parameters details are the following:  
 +  * Max Motif Error: Is the number of motif errors the user wants to allow per motif (mutations/ motif errors allowed: deletions, mismatches, insertions).  
 +  * Max adjacent ATR elements: The number of ATREs that the user allows next to each other.  
 +  * Motif Range Options: The motif range option enables the user to specify a range of motifs FireµSat should search for.  
 +  * Min required TR elements: Indicates the minimum number of TREs that should occur before a TR is output – this parameter serves as a length filter.  
 +  * Mismatch penalty (m_p): ​ The penalty value allocated to a mismatch.  
 +  * Delete penalty (d_p): ​ The penalty value assigned to a deletion.  
 +  * Insert penalty (i_p): ​ The penalty value allocated to an insertion.  
 +  * Max substring error: Is a threshold value that the user enters. The substring error should always be smaller than the Max substring error. The calculation of the substring-error (α) is simply: α = (m_p x n_m) + (i_p x n_i) + (p_d + n_d) where Mismatch count (n_m) is the number of mismatches; delete count (n_d) is the number of deletions and insert count (n_i) is the number of insertions.  
 + 
  
 ==Phobos== ​ ==Phobos== ​
Line 159: Line 169:
  
 ==QDD, QDD2== ​ ==QDD, QDD2== ​
-Meglecz et al 2010(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btp670|Meglecz,​ E., Costedoat, C., Dubut, V., Gilles, A., Malausa, T., Pech, N., and Martin, J.-F. (2009). QDD: a user-friendly program to select microsatellite markers and design primers from large sequencing projects. Bioinformatics 26, 403-404.]])) The algorithm takes the following steps: ​1) Sequence cleaning and microsatellite detection (Sort according to sequence tags, trims vector and adaptors sequences). ​2) Sequence similarity detection (Is time consuming, is meant to remove redundancy, and sequences that are part of a repetitive region of the genome). ​3) Iterative primer design using Primer3. Iterations (a) from designing primers only for perfect microsatellites with no short repeats in the flanking region till allowing multiple target microsatellites,​ short repetitions and homopolymers in the amplified region (b) produce primer pairs with PCR products in different size ranges. ​4) BLASTs selected sequences to Genbank to detect serious contaminations or mixing up samples. ​+Meglecz et al 2010(([[http://​bioinformatics.oxfordjournals.org/​cgi/​doi/​10.1093/​bioinformatics/​btp670|Meglecz,​ E., Costedoat, C., Dubut, V., Gilles, A., Malausa, T., Pech, N., and Martin, J.-F. (2009). QDD: a user-friendly program to select microsatellite markers and design primers from large sequencing projects. Bioinformatics 26, 403-404.]])) The algorithm takes the following steps: ​ 
 +  - Sequence cleaning and microsatellite detection (Sort according to sequence tags, trims vector and adaptors sequences). ​ 
 +  - Sequence similarity detection (Is time consuming, is meant to remove redundancy, and sequences that are part of a repetitive region of the genome). ​ 
 +  - Iterative primer design using Primer3. Iterations (a) from designing primers only for perfect microsatellites with no short repeats in the flanking region till allowing multiple target microsatellites,​ short repetitions and homopolymers in the amplified region (b) produce primer pairs with PCR products in different size ranges. ​ 
 +  - BLASTs selected sequences to Genbank to detect serious contaminations or mixing up samples. ​