Differences
This shows you the differences between two versions of the page.
| Both sides previous revision Previous revision Next revision | Previous revision | ||
|
bioinformatic_tools_to_detect_microsatellites_loci_from_genomic_data [2011/11/14 20:48] anniearchambault |
bioinformatic_tools_to_detect_microsatellites_loci_from_genomic_data [2013/08/08 19:21] (current) anniearchambault |
||
|---|---|---|---|
| Line 23: | Line 23: | ||
| ^Mreps | 2003 | ? | [[http://scholar.google.ca/scholar?cites=7808444266556719865&as_sdt=2005&sciodt=0,5&hl=fr|156]] | Kolpakov et al 2003(([[http://www.nar.oupjournals.org/cgi/doi/10.1093/nar/gkg617|Kolpakov, R., Bana, G., and Kucherov, G. (2003). mreps: efficient and flexible detection of tandem repeats in DNA. Nucleic Acids Research 31, 3672-3678.]])) | Sites shorter than period + 9 are automatically discarded. | Mixed combinatorial/heuristic paradigm | Start and end positions of the region to be processed; length interval, period interval, minimal exponent of the repetitions to report; resolution level, use or not of the sliding window. | One file with multiple sequences in the fasta format. | list of all repeats, with start and end positions of the repeat in the sequence, overall size of the repeat, period, exponent, error level, the repeat sequence itself | Yes, imperfect and compound (not indels) | Yes | [[http://bioinfo.lifl.fr/mreps/mreps.php|Yes]] (or [[http://mobyle.pasteur.fr/cgi-bin/portal.py?#forms::mreps|here]]) | Command-line [[http://bioinfo.lifl.fr/mreps//|Download]] | Linux, SunOS, Digital Unix and Windows systems. Online? | ANSI C | Should be fast. linear in the sequence length | ? | Automated statistical analysis files not generated; | | | ^Mreps | 2003 | ? | [[http://scholar.google.ca/scholar?cites=7808444266556719865&as_sdt=2005&sciodt=0,5&hl=fr|156]] | Kolpakov et al 2003(([[http://www.nar.oupjournals.org/cgi/doi/10.1093/nar/gkg617|Kolpakov, R., Bana, G., and Kucherov, G. (2003). mreps: efficient and flexible detection of tandem repeats in DNA. Nucleic Acids Research 31, 3672-3678.]])) | Sites shorter than period + 9 are automatically discarded. | Mixed combinatorial/heuristic paradigm | Start and end positions of the region to be processed; length interval, period interval, minimal exponent of the repetitions to report; resolution level, use or not of the sliding window. | One file with multiple sequences in the fasta format. | list of all repeats, with start and end positions of the repeat in the sequence, overall size of the repeat, period, exponent, error level, the repeat sequence itself | Yes, imperfect and compound (not indels) | Yes | [[http://bioinfo.lifl.fr/mreps/mreps.php|Yes]] (or [[http://mobyle.pasteur.fr/cgi-bin/portal.py?#forms::mreps|here]]) | Command-line [[http://bioinfo.lifl.fr/mreps//|Download]] | Linux, SunOS, Digital Unix and Windows systems. Online? | ANSI C | Should be fast. linear in the sequence length | ? | Automated statistical analysis files not generated; | | | ||
| ^STRING | 2003 | 2003 | [[http://scholar.google.ca/scholar?cites=7212951020758895252&as_sdt=2005&sciodt=0,5&hl=fr|32]] | Parisi et al 2003(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/btg268|Parisi, V., De Fonzo, V., and Aluffi-Pentini, F. (2003). STRING: finding tandem repeats in DNA sequences. Bioinformatics 19, 1733-1738.]])) | No limit in repeat length | Heuristic; dynamic programming procedure. | ? | One file with one sequence, no limit in length. | Length of the consensus word; first and last position of the TR; number of repeated units; score; consensus word; flanking sequences; alignment between the model TR and the given sequence; number of indels; number of matches and mismatches; TR base composition percentages; flag indicating a likely nested expansion. | Yes, imperfect and compound (and indels?) | No but may be possible | A web interface is mentioned in the original publication, but cannot be found | Command-line [[http://www.caspur.it/~castri/STRING/|Download]] | Unix, Windows, MacOS | C | Slowly increases as a function of the sequence length, while it increases more quickly as a function of the number of TRs. | ? | Automated statistical analysis files not generated. | | | ^STRING | 2003 | 2003 | [[http://scholar.google.ca/scholar?cites=7212951020758895252&as_sdt=2005&sciodt=0,5&hl=fr|32]] | Parisi et al 2003(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/btg268|Parisi, V., De Fonzo, V., and Aluffi-Pentini, F. (2003). STRING: finding tandem repeats in DNA sequences. Bioinformatics 19, 1733-1738.]])) | No limit in repeat length | Heuristic; dynamic programming procedure. | ? | One file with one sequence, no limit in length. | Length of the consensus word; first and last position of the TR; number of repeated units; score; consensus word; flanking sequences; alignment between the model TR and the given sequence; number of indels; number of matches and mismatches; TR base composition percentages; flag indicating a likely nested expansion. | Yes, imperfect and compound (and indels?) | No but may be possible | A web interface is mentioned in the original publication, but cannot be found | Command-line [[http://www.caspur.it/~castri/STRING/|Download]] | Unix, Windows, MacOS | C | Slowly increases as a function of the sequence length, while it increases more quickly as a function of the number of TRs. | ? | Automated statistical analysis files not generated. | | | ||
| - | ^W-SSRF | 2003 | ? | [[http://scholar.google.ca/scholar?cites=6592174860693683120&as_sdt=2005&sciodt=0,5&hl=fr|8]] | Sreenu et al 2003(([[http://www.ncbi.nlm.nih.gov/pubmed/15130803|Sreenu, V. B., Ranjitkumar, G., Swaminathan, S., Priya, S., Bose, B., Pavan, M. N., Thanu, G., Nagaraju, J., and Nagarajaram, H. A. (2003). MICAS: a fully automated web server for microsatellite extraction and analysis from prokaryote and viral genomic sequences. Appl. Bioinformatics 2, 165-168.]])) | 1 to 10 bp long | Scans a nucleotide sequence | ? | One file with one sequence. Upload limit is a 20 kb file. | Sequence content of the motif, repeat numbers, start and end position of the tract in the sequence | No, perfect only | Yes, using Autoprimer (in MICAS) | No | The user friendly GUI (graphical user interface) is MICAS, Available upon request to the [[mailto:[email protected]|authors]] | MICAS is [[http://210.212.212.7/MIC/index.html|web-only]] | Java | ? | ? | ? | | | + | ^W-SSRF | 2003 | ? | [[http://scholar.google.ca/scholar?cites=6592174860693683120&as_sdt=2005&sciodt=0,5&hl=fr|8]] | Sreenu et al 2003(([[http://www.ncbi.nlm.nih.gov/pubmed/15130803|Sreenu, V. B., Ranjitkumar, G., Swaminathan, S., Priya, S., Bose, B., Pavan, M. N., Thanu, G., Nagaraju, J., and Nagarajaram, H. A. (2003). MICAS: a fully automated web server for microsatellite extraction and analysis from prokaryote and viral genomic sequences. Appl. Bioinformatics 2, 165-168.]])) | 1 to 10 bp long | Scans a nucleotide sequence | ? | One file with one sequence. Upload limit is a 20 kb file. | Sequence content of the motif, repeat numbers, start and end position of the tract in the sequence | No, perfect only | Yes, using Autoprimer (in MICAS) | No | The user friendly GUI (graphical user interface) is MICAS, Available upon request to the [[mailto:[email protected]|authors]] | MICAS is [[http://micas.cdfd.org.in:8080/MIC/|web-only]] | Java | ? | ? | ? | | |
| ^IRF program | 2004 | 2007 | [[http://scholar.google.ca/scholar?cites=7789233353140150696&as_sdt=2005&sciodt=0,5&hl=fr|80]] | Warburton et al 2004(([[http://www.genome.org/cgi/doi/10.1101/gr.2542904|Warburton, P. E. (2004). Inverted repeat structure of the human genome: The X-chromosome contains a preponderance of large, highly homologous inverted repeats that contain testes genes. Genome Research 14, 1861-1869.]])) | ? | ? | ? | ? | ? | | | ? | Command-line [[http://tandem.bu.edu/news.html#mar2007|Download]] | ? | ? | ? | ? | ? | | | ^IRF program | 2004 | 2007 | [[http://scholar.google.ca/scholar?cites=7789233353140150696&as_sdt=2005&sciodt=0,5&hl=fr|80]] | Warburton et al 2004(([[http://www.genome.org/cgi/doi/10.1101/gr.2542904|Warburton, P. E. (2004). Inverted repeat structure of the human genome: The X-chromosome contains a preponderance of large, highly homologous inverted repeats that contain testes genes. Genome Research 14, 1861-1869.]])) | ? | ? | ? | ? | ? | | | ? | Command-line [[http://tandem.bu.edu/news.html#mar2007|Download]] | ? | ? | ? | ? | ? | | | ||
| ^ExTRS | 2004 | ? | [[http://scholar.google.ca/scholar?cites=10980601425678591553&as_sdt=2005&sciodt=0,5&hl=fr|33]] | Krishnan and Tang 2004(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bth311|Krishnan, A., and Tang, F. (2004). Exhaustive whole-genome tandem repeats search. Bioinformatics 20, 2702-2710.]])) | ? | Exhaustive | ? | One file with one sequence, no limit in length. | Redundancy in the output is reduced | Yes, substitutions (not indels) | | ? | ? Available upon request to the [[mailto:[email protected]|authors]] | ? | Source code available only on request | Near-proportional to the number of TR found | ? | ? | | | ^ExTRS | 2004 | ? | [[http://scholar.google.ca/scholar?cites=10980601425678591553&as_sdt=2005&sciodt=0,5&hl=fr|33]] | Krishnan and Tang 2004(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bth311|Krishnan, A., and Tang, F. (2004). Exhaustive whole-genome tandem repeats search. Bioinformatics 20, 2702-2710.]])) | ? | Exhaustive | ? | One file with one sequence, no limit in length. | Redundancy in the output is reduced | Yes, substitutions (not indels) | | ? | ? Available upon request to the [[mailto:[email protected]|authors]] | ? | Source code available only on request | Near-proportional to the number of TR found | ? | ? | | | ||
| Line 30: | Line 30: | ||
| ^TRA | 2004 | 2004 | [[http://scholar.google.ca/scholar?cites=13056073161710952438&as_sdt=5&sciodt=0&hl=fr|21]] | Bilgen et al 2004(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bth410|Bilgen, M., Karaca, M., Onus, A. N., and Ince, A. G. (2004). A software program combining sequence motif searches with keywords for finding repeats containing DNA sequences. Bioinformatics 20, 3379-3386.]])) | ? | Heuristic | ? | Multiple files with multiple sequences each (Max 1 Mb sequence) from ESTs. | ? | Yes, searches for exact–inexact TRs and exact–inexact compound repeats | No? | ? | ? [[ftp://ftp.akdeniz.edu.tr/Araclar/TRA/|Download]] | Windows | C++ (with Microsoft Visual C++) | ? | Searches among the organisms, organs, tissue types and development stages | ? | | | ^TRA | 2004 | 2004 | [[http://scholar.google.ca/scholar?cites=13056073161710952438&as_sdt=5&sciodt=0&hl=fr|21]] | Bilgen et al 2004(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bth410|Bilgen, M., Karaca, M., Onus, A. N., and Ince, A. G. (2004). A software program combining sequence motif searches with keywords for finding repeats containing DNA sequences. Bioinformatics 20, 3379-3386.]])) | ? | Heuristic | ? | Multiple files with multiple sequences each (Max 1 Mb sequence) from ESTs. | ? | Yes, searches for exact–inexact TRs and exact–inexact compound repeats | No? | ? | ? [[ftp://ftp.akdeniz.edu.tr/Araclar/TRA/|Download]] | Windows | C++ (with Microsoft Visual C++) | ? | Searches among the organisms, organs, tissue types and development stages | ? | | | ||
| ^MsatFinder | 2005 | 2007 | 48 | Thurston and Field 2005(([[http://www.genomics.ceh.ac.uk/msatfinder/|Thurston, M. I., and Field, D. (2006). Msatfinder, (Oxford, UK: Centre for Ecology and Hydrology. Computer program)]])) | One to 6 bp long | ? | 1) length of repeat 2) number repeat unit in the site 3) search engine (regex, multipass or iterative search) | Limit of 10 Mb of sequence in the online access. Accepts GenBank, EMBL, Swissprot, FASTA, ASCII. | Repeats, GFF, Counts, Msat_tabs, Flank_tabs, Fasta, MINE, Primers | No, but detects compound perfect repeats | | [[http://www.genomics.ceh.ac.uk/cgi-bin/msatfinder/msatfinder.cgi|Yes]] | Command-line [[http://www.genomics.ceh.ac.uk/msatfinder/#download|Download]] | Unix (may work on Mac OSX) | perl script | ? | Nucleic acid or amino acid sequence | ? | | | ^MsatFinder | 2005 | 2007 | 48 | Thurston and Field 2005(([[http://www.genomics.ceh.ac.uk/msatfinder/|Thurston, M. I., and Field, D. (2006). Msatfinder, (Oxford, UK: Centre for Ecology and Hydrology. Computer program)]])) | One to 6 bp long | ? | 1) length of repeat 2) number repeat unit in the site 3) search engine (regex, multipass or iterative search) | Limit of 10 Mb of sequence in the online access. Accepts GenBank, EMBL, Swissprot, FASTA, ASCII. | Repeats, GFF, Counts, Msat_tabs, Flank_tabs, Fasta, MINE, Primers | No, but detects compound perfect repeats | | [[http://www.genomics.ceh.ac.uk/cgi-bin/msatfinder/msatfinder.cgi|Yes]] | Command-line [[http://www.genomics.ceh.ac.uk/msatfinder/#download|Download]] | Unix (may work on Mac OSX) | perl script | ? | Nucleic acid or amino acid sequence | ? | | | ||
| - | ^FireµSat | 2006 | 2011 | [[http://scholar.google.ca/scholar?cites=17060705618169575807&as_sdt=2005&sciodt=0,5&hl=fr|5]] and [[http://scholar.google.ca/scholar?cites=66353776030260478&as_sdt=5&sciodt=0&hl=fr|this]] | de Ridder at al 2006(([[http://portal.acm.org/citation.cfm?doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). FireµSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing countries - SAICSIT ’06 (Somerset West, South Africa), pp. 247-256.]])) and [[http://upetd.up.ac.za/thesis/available/etd-08172010-202532/unrestricted/dissertation.pdf|de Ridder 2006]] | 1 to 5 bp. Length set by the user. The next update should allow for detection of 6 to 100 bp repeats. | Uses Counting Finite Automata (which are regular language acceptors) | Max Motif Error (per motif); Max adjacent ATR elements; Motif Range Options; Min required TR elements; Max substring error (a threshold); Mismatch penalty (m_p); Delete penalty (d_p); Insert penalty (i_p). | One fasta file with one sequence | File in .csv format. | Yes, substitutions and indels, but not compound loci. | No | No | GUI and command-line, [[http://www.dna-algo.co.za/|Download]] | Windows; Linux in progress | C++ and MatLab | Run time increases linearly with the sequence length; does not increase with longer motif lengths. | Designed for microsatellites, but can detect any type of TR | Fast, simple and flexible. | | | + | ^FireµSat | 2006 | 2011 | [[http://scholar.google.ca/scholar?cites=17060705618169575807&as_sdt=2005&sciodt=0,5&hl=fr|5]] and [[http://scholar.google.ca/scholar?cites=66353776030260478&as_sdt=5&sciodt=0&hl=fr|this]] | de Ridder at al 2006(([[http://portal.acm.org/citation.cfm?doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). FireµSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing countries - SAICSIT ’06 (Somerset West, South Africa), pp. 247-256.]])) and de Ridder at al 2013(([[http://www.sciencedirect.com/science/article/pii/S1570866712001657|De Ridder, C., D.G. Kourie, B.W. Watson, T.R. Fourie, and P.V. Reyneke (2013). Fine-tuning the search for microsatellites. Journal of Discrete Algorithms 20: 21–37.]])) | 1 to 5 bp. Length set by the user. The next update should allow for detection of 6 to 100 bp repeats. | Uses Counting Finite Automata (which are regular language acceptors) | Max Motif Error (per motif); Max adjacent ATR elements; Motif Range Options; Min required TR elements; Max substring error (a threshold); Mismatch penalty (m_p); Delete penalty (d_p); Insert penalty (i_p). | One fasta file with one sequence | File in .csv format. | Yes, substitutions and indels, but not compound loci. | No | No | GUI and command-line, [[http://www.dna-algo.co.za/|Download]] | Windows; Linux in progress | C++ and MatLab | Run time increases linearly with the sequence length; does not increase with longer motif lengths. | Designed for microsatellites, but can detect any type of TR | Fast, simple and flexible. | | |
| ^Phobos | 2006 | 2010 | ? | Mayer 2010(([[http://www.ruhr-uni-bochum.de/spezzoo/cm/cm_phobos.htm |Mayer, C. (2010). Phobos: Highly accurate search for perfect and imperfect tandem repeats in complete genomes by Christoph Mayer, (Bochum, Germany: Ruhr-Universität Bochum,Faculty of Biological Sciences and Biotechnology). Computer program.]])) | Perfect and imperfect TR, with a pattern size of 1 - 10 000 bp | Exhaustive, uses alignment scores | Mismatch score, indel score, minimum score, minimum length, minimum perfection, and others. | One file in fasta format, with multiple sequence. No limit in sequence length. | Text file, different formats, including gff and fasta | Yes. Substitutions and indels. | Not in itself, but yes as implemented in STAMP or Geneious. | No | User friendly GUI and easily scriptable Command-line program [[http://www.ruhr-uni-bochum.de/ecoevo/cm/cm_phobos.htm|Download]] | MacOSX, Linux, Windows. | C++ | Execution time increases with pattern size range. Very fast in the size range 1-10 bp, slow for patterns in the size range above 10-20 bp. | Can be incorporated into pipelines. Implemented in STAMP and Geneious. | Free only for academic users. | | | ^Phobos | 2006 | 2010 | ? | Mayer 2010(([[http://www.ruhr-uni-bochum.de/spezzoo/cm/cm_phobos.htm |Mayer, C. (2010). Phobos: Highly accurate search for perfect and imperfect tandem repeats in complete genomes by Christoph Mayer, (Bochum, Germany: Ruhr-Universität Bochum,Faculty of Biological Sciences and Biotechnology). Computer program.]])) | Perfect and imperfect TR, with a pattern size of 1 - 10 000 bp | Exhaustive, uses alignment scores | Mismatch score, indel score, minimum score, minimum length, minimum perfection, and others. | One file in fasta format, with multiple sequence. No limit in sequence length. | Text file, different formats, including gff and fasta | Yes. Substitutions and indels. | Not in itself, but yes as implemented in STAMP or Geneious. | No | User friendly GUI and easily scriptable Command-line program [[http://www.ruhr-uni-bochum.de/ecoevo/cm/cm_phobos.htm|Download]] | MacOSX, Linux, Windows. | C++ | Execution time increases with pattern size range. Very fast in the size range 1-10 bp, slow for patterns in the size range above 10-20 bp. | Can be incorporated into pipelines. Implemented in STAMP and Geneious. | Free only for academic users. | | | ||
| ^SSRscanner | 2006 | ? | [[http://scholar.google.ca/scholar?cites=15246530144073441813&as_sdt=2005&sciodt=0,5&hl=fr|3]] | Anwar and Khan 2006(([[http://www.ncbi.nlm.nih.gov/pmc/articles/PMC1891659/|Anwar, T., and Khan, A. (2006). SSRscanner: a program for reporting distribution and exact location of simple sequence repeats. Bioinformation 1, 89-91.]])) | Only searches for predefined motifs | Exhaustive, uses dictionary approach. | File containing motifs of different repeat types; number of times for the motifs to be repeated | One file with one sequence | Motifposition.txt (gives the frequency of each repeat provided in the motif file) and (2) Motifresult.exe (gives the specific location of each repeat) | No, perfect only | No | No | Command-line Availability unknown, contact the [[mailto:[email protected]|author]] | Platform independent | perl script | ? | ? | ? | | | ^SSRscanner | 2006 | ? | [[http://scholar.google.ca/scholar?cites=15246530144073441813&as_sdt=2005&sciodt=0,5&hl=fr|3]] | Anwar and Khan 2006(([[http://www.ncbi.nlm.nih.gov/pmc/articles/PMC1891659/|Anwar, T., and Khan, A. (2006). SSRscanner: a program for reporting distribution and exact location of simple sequence repeats. Bioinformation 1, 89-91.]])) | Only searches for predefined motifs | Exhaustive, uses dictionary approach. | File containing motifs of different repeat types; number of times for the motifs to be repeated | One file with one sequence | Motifposition.txt (gives the frequency of each repeat provided in the motif file) and (2) Motifresult.exe (gives the specific location of each repeat) | No, perfect only | No | No | Command-line Availability unknown, contact the [[mailto:[email protected]|author]] | Platform independent | perl script | ? | ? | ? | | | ||
| Line 38: | Line 38: | ||
| ^SciRoKo | 2007 | 2008 | [[http://scholar.google.ca/scholar?cites=16996584693111364287&as_sdt=2005&sciodt=0,5&hl=fr|44]] | Kofler et al 2007(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/btm157|Kofler, R., Schlotterer, C., and Lelley, T. (2007). SciRoKo: a new tool for whole genome microsatellite search and investigation. Bioinformatics 23, 1683-1685.]])) | One to 6 bp long | ? | Hits (identity with a virtual perfect microsatellite), number of mismatches (mm), mismatch penalty (mmP) and the length of the SSR motif (mL). | One file with multiple sequences in the fasta format. | ? | Yes, imperfect and compound (indels?) | | No | User friendly GUI (graphical user interface) standalone. [[http://www.kofler.or.at/bioinformatics/SciRoKo/index.html|Download]] | Windows. Should be platform independent, but Mac users have not been able install | C# | Fast | ? | Depends on .NET framework | | | ^SciRoKo | 2007 | 2008 | [[http://scholar.google.ca/scholar?cites=16996584693111364287&as_sdt=2005&sciodt=0,5&hl=fr|44]] | Kofler et al 2007(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/btm157|Kofler, R., Schlotterer, C., and Lelley, T. (2007). SciRoKo: a new tool for whole genome microsatellite search and investigation. Bioinformatics 23, 1683-1685.]])) | One to 6 bp long | ? | Hits (identity with a virtual perfect microsatellite), number of mismatches (mm), mismatch penalty (mmP) and the length of the SSR motif (mL). | One file with multiple sequences in the fasta format. | ? | Yes, imperfect and compound (indels?) | | No | User friendly GUI (graphical user interface) standalone. [[http://www.kofler.or.at/bioinformatics/SciRoKo/index.html|Download]] | Windows. Should be platform independent, but Mac users have not been able install | C# | Fast | ? | Depends on .NET framework | | | ||
| ^Msatcommander | 2008 | 2011 | [[http://scholar.google.ca/scholar?cites=3185583603820687905&as_sdt=2005&sciodt=0,5&hl=fr|102]] | Faircloth 2008(([[http://doi.wiley.com/10.1111/j.1471-8286.2007.01884.x|Faircloth, B. C. (2008). msatcommander: detection of microsatellite repeat arrays and automated, locus-specific primer design. Molecular Ecology Resources 8, 92-94.]])) | ? | Uses regular expressions | ? | One file with multiple sequences in the fasta format. | Either a summary file (array detection only) or a directory at a user-selectable location | No, but accepts N | Yes, with Primer3, includes 5'-tailing | No | User friendly GUI (graphical user interface) standalone. [[http://code.google.com/p/msatcommander/|Download]] | MacOS X, Windows, Unix. | Python | ? | Rapid and automated microsatellite array detection, locus-specific primer design, and 5'-tailing of designed primers | ? | | | ^Msatcommander | 2008 | 2011 | [[http://scholar.google.ca/scholar?cites=3185583603820687905&as_sdt=2005&sciodt=0,5&hl=fr|102]] | Faircloth 2008(([[http://doi.wiley.com/10.1111/j.1471-8286.2007.01884.x|Faircloth, B. C. (2008). msatcommander: detection of microsatellite repeat arrays and automated, locus-specific primer design. Molecular Ecology Resources 8, 92-94.]])) | ? | Uses regular expressions | ? | One file with multiple sequences in the fasta format. | Either a summary file (array detection only) or a directory at a user-selectable location | No, but accepts N | Yes, with Primer3, includes 5'-tailing | No | User friendly GUI (graphical user interface) standalone. [[http://code.google.com/p/msatcommander/|Download]] | MacOS X, Windows, Unix. | Python | ? | Rapid and automated microsatellite array detection, locus-specific primer design, and 5'-tailing of designed primers | ? | | | ||
| - | ^ReRep | 2008 | 2008 | [[http://scholar.google.ca/scholar?cites=255635357253003447&as_sdt=2005&sciodt=0,5&hl=fr|5]] | Otto et al 2008(([[http://www.biomedcentral.com/1471-2105/9/366|Otto, T. D., Gomes, L. H. F., Alves-Ferreira, M., de Miranda, A. B., and Degrave, W. M. (2008). ReRep: Computational detection of repetitive sequences in genome survey sequences (GSS). BMC Bioinformatics 9, 366.]])) | ? | Uses self-similarity searches | ? | Genome survey sequences(GSS) files, including 454-reads | ? | Yes, substitutions and indels | No | No | Command-line [[http://bioinfo.pdtis.fiocruz.br/ReRep/|Download]] | Linux | Perl | ? | Can detect de novo repeats in Genome Sequence Survey sequence data | ? | | + | ^ReRep | 2008 | 2008 | [[http://scholar.google.ca/scholar?cites=255635357253003447&as_sdt=2005&sciodt=0,5&hl=fr|5]] | Otto et al 2008(([[http://www.biomedcentral.com/1471-2105/9/366|Otto, T. D., Gomes, L. H. F., Alves-Ferreira, M., de Miranda, A. B., and Degrave, W. M. (2008). ReRep: Computational detection of repetitive sequences in genome survey sequences (GSS). BMC Bioinformatics 9, 366.]])) | ? | Uses self-similarity searches | ? | Genome survey sequences(GSS) files, including 454-reads | ? | Yes, substitutions and indels | No | No | Command-line [[http://www.dbbm.fiocruz.br/labwim/bioinfoteam/index.pl?action=services|Download]] | Linux | Perl | ? | Can detect de novo repeats in Genome Sequence Survey sequence data | ? | |
| ^T-REKS | 2009 | ? | [[http://scholar.google.ca/scholar?cites=208721574341442296&as_sdt=2005&sciodt=0,5&hl=fr|5]] | Jorda and Kajava 2009(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/btp482|Jorda, J., and Kajava, A. V. (2009). T-REKS: identification of Tandem REpeats in sequences with a K-meanS based algorithm. Bioinformatics 25, 2632-2638.]])) | No limits in repeat length | Short string extension and K-means algorithm. | delta-l Allowed % of length variability, P*sim—similarity threshold and an option to allow or not the detection of overlaping TR. | Sequences in FASTA format. | Output with start, end, length of TR and multiple alignment of the repeats. | Yes, substitutions and indels. | No | [[http://bioinfo.montp.cnrs.fr/?r=t-reks|Yes]] | User friendly GUI (graphical user interface) standalone. [[http://bioinfo.montp.cnrs.fr/?r=t-reks|Download]] | Platform independent | Java | Fast. Execution time is linear (directly proportional) to the sequence length. | Can be applied to nucleic acid, amino acid or any text sequence. | ? | | ^T-REKS | 2009 | ? | [[http://scholar.google.ca/scholar?cites=208721574341442296&as_sdt=2005&sciodt=0,5&hl=fr|5]] | Jorda and Kajava 2009(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/btp482|Jorda, J., and Kajava, A. V. (2009). T-REKS: identification of Tandem REpeats in sequences with a K-meanS based algorithm. Bioinformatics 25, 2632-2638.]])) | No limits in repeat length | Short string extension and K-means algorithm. | delta-l Allowed % of length variability, P*sim—similarity threshold and an option to allow or not the detection of overlaping TR. | Sequences in FASTA format. | Output with start, end, length of TR and multiple alignment of the repeats. | Yes, substitutions and indels. | No | [[http://bioinfo.montp.cnrs.fr/?r=t-reks|Yes]] | User friendly GUI (graphical user interface) standalone. [[http://bioinfo.montp.cnrs.fr/?r=t-reks|Download]] | Platform independent | Java | Fast. Execution time is linear (directly proportional) to the sequence length. | Can be applied to nucleic acid, amino acid or any text sequence. | ? | | ||
| ^BwTRS | 2010 | 2009 | [[http://scholar.google.ca/scholar?cites=7603849771779688492&as_sdt=2005&sciodt=0,5&hl=fr|2]] | Pokrzywa and Polanski 2010(([[http://linkinghub.elsevier.com/retrieve/pii/S0888754310001758|Pokrzywa, R., and Polanski, A. (2010). BWtrs: A tool for searching for tandem repeats in DNA sequences based on the Burrows–Wheeler transform. Genomics 96, 316-321.]])) | ? | Exhaustive; uses efficient data compression algorithm | "Minimum motif size", "Maximum motif size", "Minimum repeat size" and "Minimum repeat ratio". | One file with multiple sequences in the fasta format; or GenBank id. Accepts nucleotides and amino acids sequences. | List of all TR with: Start and End (position of the TR in the sequence); Motif length; Ratio (between the motif length and the consensus repeat length); the motif itself. HTML or text. | No | No | [[http://bwtools.polsl.pl/BWtrs/input.jsp|Yes]], runs on a standard Tomcat servlet container without any local database | ? Availability unknown, contact the [[mailto: [email protected]|authors]]. | ? | Java | Depends on sequence length | Nucleic acid or amino acid sequence | ? | | ^BwTRS | 2010 | 2009 | [[http://scholar.google.ca/scholar?cites=7603849771779688492&as_sdt=2005&sciodt=0,5&hl=fr|2]] | Pokrzywa and Polanski 2010(([[http://linkinghub.elsevier.com/retrieve/pii/S0888754310001758|Pokrzywa, R., and Polanski, A. (2010). BWtrs: A tool for searching for tandem repeats in DNA sequences based on the Burrows–Wheeler transform. Genomics 96, 316-321.]])) | ? | Exhaustive; uses efficient data compression algorithm | "Minimum motif size", "Maximum motif size", "Minimum repeat size" and "Minimum repeat ratio". | One file with multiple sequences in the fasta format; or GenBank id. Accepts nucleotides and amino acids sequences. | List of all TR with: Start and End (position of the TR in the sequence); Motif length; Ratio (between the motif length and the consensus repeat length); the motif itself. HTML or text. | No | No | [[http://bwtools.polsl.pl/BWtrs/input.jsp|Yes]], runs on a standard Tomcat servlet container without any local database | ? Availability unknown, contact the [[mailto: [email protected]|authors]]. | ? | Java | Depends on sequence length | Nucleic acid or amino acid sequence | ? | | ||
| Line 46: | Line 46: | ||
| ^repeatfinder | 2001 | ? | ? | Volfovsky et al 2001(([[http://genomebiology.com/2001/2/8/research/0027|Volfovsky, N., Haas, B., and Salzberg, S. (2001). A clustering method for repeat analysis in DNA sequences. Genome Biology 2, research0027.1 - research0027.11.]])) | ? | Uses REPuter | ? | ? | ? | ? | ? | No | Command-line [[http://www.cbcb.umd.edu/software/RepeatFinder/|Download]] | Linux RedHat 6.x+, Sun Solaris, and Alpha OSF1 | ? Open Source | ? | ? | ? | | ^repeatfinder | 2001 | ? | ? | Volfovsky et al 2001(([[http://genomebiology.com/2001/2/8/research/0027|Volfovsky, N., Haas, B., and Salzberg, S. (2001). A clustering method for repeat analysis in DNA sequences. Genome Biology 2, research0027.1 - research0027.11.]])) | ? | Uses REPuter | ? | ? | ? | ? | ? | No | Command-line [[http://www.cbcb.umd.edu/software/RepeatFinder/|Download]] | Linux RedHat 6.x+, Sun Solaris, and Alpha OSF1 | ? Open Source | ? | ? | ? | | ||
| ^MsatMiner | 2005 | ? | | Thurston and Field 2005(([[http://www.genomics.ceh.ac.uk/msatminer/|Thurston, M. I. (2005). Msatminer - scripts for processing msatfinder output, (Oxford, UK: Centre for Ecology and Hydrology). Computer program.]])) | ? | Uses msatfinder for motif discovery | | Different format of sequence file | More statistical analysis are possible. | Yes, imperfect and compound (indels?) | Possible when using additional scripts | ? | Command-line | Unix and MacOS | Collection of perl scripts | ? | ? | Running scripts is possibly complicated | | ^MsatMiner | 2005 | ? | | Thurston and Field 2005(([[http://www.genomics.ceh.ac.uk/msatminer/|Thurston, M. I. (2005). Msatminer - scripts for processing msatfinder output, (Oxford, UK: Centre for Ecology and Hydrology). Computer program.]])) | ? | Uses msatfinder for motif discovery | | Different format of sequence file | More statistical analysis are possible. | Yes, imperfect and compound (indels?) | Possible when using additional scripts | ? | Command-line | Unix and MacOS | Collection of perl scripts | ? | ? | Running scripts is possibly complicated | | ||
| - | ^E-TRA | 2005 | 2004 | [[http://scholar.google.ca/scholar?cites=6055835177193667419&as_sdt=5&sciodt=0&hl=fr|12]] | Karaca et al 2005(([[Available at: http://www.springerlink.com/index/10.1007/BF02715889|Karaca, M., Bilgen, M., Onus, A. N., Ince, A. G., and Elmasulu, S. Y. (2005). Exact tandem repeats analyzer (E-TRA): A new program for DNA sequence mining. J Genet 84, 49-54.]])) | 1 to 1000 bp repeat | Uses TRA | ? | Multiple files with multiple sequences each (maximum of 1 Mb long) | | Yes, compound and imperfect | Yes | No | user friendly GUI (graphical user interface) [[ftp://ftp.akdeniz.edu.tr/Araclar/TRA/|Download]] | Windows | C++ (with Microsoft Visual C++) | | Searches among the organisms, organs, tissue types and development stages | Only 1 Mb of input sequence length | | + | ^E-TRA | 2005 | 2004 | [[http://scholar.google.ca/scholar?cites=6055835177193667419&as_sdt=5&sciodt=0&hl=fr|12]] | Karaca et al 2005(([[http://www.springerlink.com/index/10.1007/BF02715889|Karaca, M., Bilgen, M., Onus, A. N., Ince, A. G., and Elmasulu, S. Y. (2005). Exact tandem repeats analyzer (E-TRA): A new program for DNA sequence mining. J Genet 84, 49-54.]])) | 1 to 1000 bp repeat | Uses TRA | ? | Multiple files with multiple sequences each (maximum of 1 Mb long) | | Yes, compound and imperfect | Yes | No | user friendly GUI (graphical user interface) [[ftp://ftp.akdeniz.edu.tr/Araclar/TRA/|Download]] | Windows | C++ (with Microsoft Visual C++) | | Searches among the organisms, organs, tissue types and development stages | Only 1 Mb of input sequence length | |
| - | ^SSRprimerII | 2006 | 2009 | 29 | Robinson et al. 2004(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bth104|Robinson, A. J., Love, C. G., Batley, J., Barker, G., and Edwards, D. (2004). Simple sequence repeat marker loci discovery using SSR primer. Bioinformatics 20, 1475-1476.]])) and Jewell et al. 2006(([[http://www.nar.oxfordjournals.org/cgi/doi/10.1093/nar/gkl083|Jewell, E. et al. (2006). SSRPrimer and SSR Taxonomy Tree: Biome SSR discovery. Nucleic Acids Research 34, W656-W659.]])) | 2 to >6 bp long repeats | Uses Sputnik | | Limit of 4000 bp sequence | | | Yes uses Primer3 | [[http://flora.acpfg.com.au/ssrprimer2/cgi-bin/index|Yes]] | ? Contact the [[ttp://www.appliedbioinformatics.com.au/|group]] | ? | Perl scripts | ? | ? | ? | | + | ^SSRprimerII | 2006 | 2009 | 29 | Robinson et al. 2004(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bth104|Robinson, A. J., Love, C. G., Batley, J., Barker, G., and Edwards, D. (2004). Simple sequence repeat marker loci discovery using SSR primer. Bioinformatics 20, 1475-1476.]])) and Jewell et al. 2006(([[http://www.nar.oxfordjournals.org/cgi/doi/10.1093/nar/gkl083|Jewell, E. et al. (2006). SSRPrimer and SSR Taxonomy Tree: Biome SSR discovery. Nucleic Acids Research 34, W656-W659.]])) | 2 to >6 bp long repeats | Uses Sputnik | | Limit of 4000 bp sequence | | | Yes uses Primer3 | [[http://flora.acpfg.com.au/ssrprimer2/cgi-bin/index|Yes]] | ? Contact the [[http://www.appliedbioinformatics.com.au/|group]] | ? | Perl scripts | ? | ? | ? | |
| ^TRAP | 2006 | 2005 | 10 | Sobreira et al 2006(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bti809|Sobreira, T. J. P., Durham, A. M., and Gruber, A. (2005). TRAP: automated classification, quantification and annotation of tandemly repeated sequences. Bioinformatics 22, 361-362.]])) | 1 to 2000 bp repeat | Uses TRF | All TRF parameters, as well as: min and max number of motifs; min and max motif size (period), min size of flanking regions and min match % between adjacent motifs. | One file with multiple sequences in fasta format. | Many formats (csv, HTML, flat files, and GFF) | Yes, imperfect and compound | No, but generates a fasta file with TR regions masked with Ns | no | Command-line [[http://www.coccidia.icb.usp.br/trap/trapRegister/|Download]] | Unix; MacOSX | Perl scripts | Near-proportional to the number of TRs found by TRF | Selection, classification, quantification and automated annotation of TR sequences | Not compatible with MS Windows version of TRF | | ^TRAP | 2006 | 2005 | 10 | Sobreira et al 2006(([[http://bioinformatics.oxfordjournals.org/cgi/doi/10.1093/bioinformatics/bti809|Sobreira, T. J. P., Durham, A. M., and Gruber, A. (2005). TRAP: automated classification, quantification and annotation of tandemly repeated sequences. Bioinformatics 22, 361-362.]])) | 1 to 2000 bp repeat | Uses TRF | All TRF parameters, as well as: min and max number of motifs; min and max motif size (period), min size of flanking regions and min match % between adjacent motifs. | One file with multiple sequences in fasta format. | Many formats (csv, HTML, flat files, and GFF) | Yes, imperfect and compound | No, but generates a fasta file with TR regions masked with Ns | no | Command-line [[http://www.coccidia.icb.usp.br/trap/trapRegister/|Download]] | Unix; MacOSX | Perl scripts | Near-proportional to the number of TRs found by TRF | Selection, classification, quantification and automated annotation of TR sequences | Not compatible with MS Windows version of TRF | | ||
| ^cid | 2008 | ? | 11 | Freita et al 2008(([[http://doi.wiley.com/10.1111/j.1471-8286.2007.01950.x|Freitas, P. D., Martins, D. S., and Galetti, P. M. (2008). cid: a rapid and efficient bioinformatic tool for the detection of SSRs from genomic libraries. Molecular Ecology Resources 8, 107-108.]])) | Same as MISA | Uses MISA for tandem repeat detection, and other external programs for other steps | | Set of chromatograms or multiFASTA file | List of useful primers | | Yes, using Primer3 | ? | Web environment, Availability unknown, contact the [[mailto: [email protected]|authors]] | ? | perl and php to connect the different tools | ? | Can mask vectors and adaptors regions of cloned sequences | ? | | ^cid | 2008 | ? | 11 | Freita et al 2008(([[http://doi.wiley.com/10.1111/j.1471-8286.2007.01950.x|Freitas, P. D., Martins, D. S., and Galetti, P. M. (2008). cid: a rapid and efficient bioinformatic tool for the detection of SSRs from genomic libraries. Molecular Ecology Resources 8, 107-108.]])) | Same as MISA | Uses MISA for tandem repeat detection, and other external programs for other steps | | Set of chromatograms or multiFASTA file | List of useful primers | | Yes, using Primer3 | ? | Web environment, Availability unknown, contact the [[mailto: [email protected]|authors]] | ? | perl and php to connect the different tools | ? | Can mask vectors and adaptors regions of cloned sequences | ? | | ||
| Line 126: | Line 126: | ||
| ==FireµSat== | ==FireµSat== | ||
| - | de Ridder at al 2006(([[http://portal.acm.org/citation.cfm?doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). FireµSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing couuntries - SAICSIT ’06 (Somerset West, South Africa), pp. 247-256.]])) Is a combination of straightforward FA technology combined with a flavour of Moore machine technology. It uses Counting Finite Automata, which are regular language acceptors. The parameters details are the following: | + | de Ridder at al 2006(([[http://portal.acm.org/citation.cfm?doid=1216262.1216289|de Riddera, C., Kourie, D. G., and Watson, B. W. (2006). FireµSat. In Proceedings of the 2006 annual research conference of the South African institute of computer scientists and information technologists on IT research in developing couuntries - SAICSIT ’06 (Somerset West, South Africa), pp. 247-256.]])) and de Ridder at al 2013(([[http://www.sciencedirect.com/science/article/pii/S1570866712001657|De Ridder, C., D.G. Kourie, B.W. Watson, T.R. Fourie, and P.V. Reyneke (2013). Fine-tuning the search for microsatellites. Journal of Discrete Algorithms 20: 21–37.]])), Is a combination of straightforward FA technology combined with a flavour of Moore machine technology. It uses Counting Finite Automata, which are regular language acceptors. The parameters details are the following: |
| * Max Motif Error: Is the number of motif errors the user wants to allow per motif (mutations/ motif errors allowed: deletions, mismatches, insertions). | * Max Motif Error: Is the number of motif errors the user wants to allow per motif (mutations/ motif errors allowed: deletions, mismatches, insertions). | ||
| * Max adjacent ATR elements: The number of ATREs that the user allows next to each other. | * Max adjacent ATR elements: The number of ATREs that the user allows next to each other. | ||
