Hello Greengenes pro!Please note, this is listed as a "term" position. According to the position description, you'll be hired for one year. Your term position may be extended after that or converted into a career position depending on your performance and future funding.
I'm including you in this job announcement since you have used multiple tools on the Greengenes 16S rRNA analysis web site. I'm guessing you are likely aware that Greengenes is both a DNA database and a web service and we've been funded to grow in both areas. If you know a recent Bachelors-level graduate with solid programming skills whose looking for an entry level full-time position, please forward them this link to the job description:
http://jobs.lbl.gov/LBNLCareers/details.asp?jid=23315&p=1sid=2027
Thanks for considering this opportunity and for all your helpful suggestions over the years.
Showing posts with label bioinformatics. Show all posts
Showing posts with label bioinformatics. Show all posts
Wednesday, August 19, 2009
Job Announcement
Received this from the people who develop Greengenes. If you're just beginning a career in bioinformatics, this sounds like a good place to start.
Friday, April 17, 2009
Linkage ...
... to some web-based resources I've had to use lately.
Libshuff (Sequence Library Comparison) - To determine whether or not your various DNA (in my case microbial) libraries are statistically similar or not. It's hosted by UGA.
DNADIST - In order to use the web-based program above, you need a DNA distance matrix file. It's a part of the Phylip package. There is an online resource for this though hosted by the Pasteur Institute (which is what is linked above).
Or, if you don't want to use libshuff like I am, you can use UniFrac.
Libshuff (Sequence Library Comparison) - To determine whether or not your various DNA (in my case microbial) libraries are statistically similar or not. It's hosted by UGA.
DNADIST - In order to use the web-based program above, you need a DNA distance matrix file. It's a part of the Phylip package. There is an online resource for this though hosted by the Pasteur Institute (which is what is linked above).
Or, if you don't want to use libshuff like I am, you can use UniFrac.
Friday, December 05, 2008
Hello Ubuntu
The new computer arrived, but in the meantime I've been given a to-be-scrapped computer to use as a testbed for my "linux project". What I'm hoping to do is build a linux cluster. At first it'll just be two computers, but I hope to expand it as things go to surplus.
Why? Because with the amount of information we'll be acquiring (especially once we start 454 sequencing in earnest), I'll need the computing power for processing the data. I'd like a 64 bit computer, but from a Windows standpoint that means using Vista, and the government doesn't currently allow Vista on their computers without seeking exceptions ... and that paperwork is a bit of a pain. Besides, several of my programs don't work on Vista yet, which means I'd need to keep an XP machine around anyways (rendering the above moot). *sigh*
Anyways, if I cluster a couple of linux machines, I can achieve (I believe) similar results. So, anyone build a cluster recently? Also, what bioinformatic tools do you use regularly on linux? I've already installed Geneious and the Bio-Linux base programs (which is mostly Emboss) and Artemis.
Why? Because with the amount of information we'll be acquiring (especially once we start 454 sequencing in earnest), I'll need the computing power for processing the data. I'd like a 64 bit computer, but from a Windows standpoint that means using Vista, and the government doesn't currently allow Vista on their computers without seeking exceptions ... and that paperwork is a bit of a pain. Besides, several of my programs don't work on Vista yet, which means I'd need to keep an XP machine around anyways (rendering the above moot). *sigh*
Anyways, if I cluster a couple of linux machines, I can achieve (I believe) similar results. So, anyone build a cluster recently? Also, what bioinformatic tools do you use regularly on linux? I've already installed Geneious and the Bio-Linux base programs (which is mostly Emboss) and Artemis.
Tuesday, November 11, 2008
Bioinformatics
So, for the microbiologists/molecular biologists/geneticists who read this blog (I know there are a couple), what programs/software do you use for sequence analysis/protein analysis/molecular biological applications?
Here's my list:
IN HOUSE
Geneious ($249 subscription/year)
Used: Daily
Official description:
Geneious Pro is an integrated, cross-platform bioinformatics software suite for manipulating, finding, sharing, and exploring biological data such as DNA sequences or proteins, phylogenies, 3D structure information, publications, etc. It features sequence alignment and phylogenetic analysis, contig assembly, primer design and restriction analysis, access to NCBI and UniProt, BLAST, protein structure viewing, automated PubMed searching, and more. It even includes an API for creating your own plugins.
What I use it for:
Geneious is the workhorse application for DNA sequence analysis (chromatogram/sequence quality) and editing (vector and quality trimming) in my laboratory. Geneious is also used for contig assembly of genes/organisms and for alignment of 16S sequences for downstream phylogeny analysis (see programs MEGA, DnaSP, DAMBE). It can also be used to construct quick phylogenetic trees for routine examination. The subscription package allows me to receive regular updates. The only other comparable software application that I’ve found that works well on Windows XP is Sequencher (2007 quote for purchase was $2975. Major updates would require another purchase).
Artemis (freeware)
Used: Moderately (several times a month)
Official description:
Artemis is a free genome viewer and annotation tool that allows visualization of sequence features and the results of analyses within the context of the sequence, and its six-frame translation. Artemis is written in Java, and is available for UNIX, GNU/Linux, BSD, Macintosh and MS Windows systems. It can read complete EMBL and GENBANK database entries or sequence in FASTA or raw format. Extra sequence features can be in EMBL, GENBANK or GFF format.
What I use it for:
Artemis is a valuable tool for examining completed genomes. Search by gene/sequence/functional category for items of interest. GenBank houses over 630 completed microbial genomes (631 as of 02/08/08).
MEGA ver4.0 (freeware)
Used: Moderately
Official description:
MEGA is an integrated tool for conducting automatic and manual sequence alignment, inferring phylogenetic trees, mining web-based databases, estimating rates of molecular evolution, and testing evolutionary hypotheses.
What I use it for:
MEGA is the primary phylogenetic tree building program. It constructs publication quality phylogenetic trees. It is used for molecular evolution and population genetic analysis. In terms of alignment data, MEGA is an established format and most programs export/import alignments in MEGA format. Geneious exports alignment data in MEGA format, allowing these two programs to be used in conjunction. MEGA also has a sequence editor for quick/minor alignment editing.
DnaSP (freeware)
Used: Infrequently/Rarely
Official description:
DnaSP, DNA Sequence Polymorphism, is a software package for the analysis of nucleotide polymorphism from aligned DNA sequence data. DnaSP can estimate several measures of DNA sequence variation within and between populations (in noncoding, synonymous or nonsynonymous sites, or in various sorts of codon positions), as well as linkage disequilibrium, recombination, gene flow and gene conversion parameters. DnaSP can also carry out several tests of neutrality: Hudson, Kreitman and Aguadé, Tajima, McDonald and Kreitman, Fu and Li, and Fu tests. Additionally, DnaSP can estimate the confidence intervals of some test-statistics by the coalescent. The results of the analyses are displayed on tabular and graphic form.
What I use it for:
DnaSP is primarily used to determine genotype numbers. Genotypes are based on SNP information derived from DNA sequencing (typically MLST – multi-locus sequence typing) closely related strains/isolates.
DAMBE (freeware)
Used: Infrequently/Rarely
Official description:
Data analysis in molecular biology and evolution. t is an integrated software package for retrieving, organizing, manipulating, aligning, and analyzing molecular sequence data. Allele frequency data can also be used by DAMBE for calculating genetic distances or phylogenetic reconstruction.
What I use it for:
DAMBE does not see frequent usage in the lab, but it is sometimes useful for determining genotype numbers (it ignores gapped sequences for example) from complex sequences.
TotalLab 120 DM ($6,000)
Used: Moderately
Official description:
The TL120 version in the TotalLab range is an advanced image analysis solution which offers an extensive range of features for the in-depth analysis of 1D electrophoresis gels and performing band pattern matching studies. TL120 DM is the TL120 analysis software complete with the DM database component so you can archive all your analysed results and perform cross experiment investigations.
What I use it for:
This program is integral for analysis of our ribosomal intergenic spacer analysis (RISA) data which is collected on a LiCor DNA sequencer. This allows us to look at archaea, eubacterial and fungal population patterns in a sample and then compare that gel image to other images/samples to construct a phylogenetic relationship between them. The DM option allows us to store this information in a database for comparison of data between experiments. This will enhance our ability to compare samples across time (date of analysis & time of collection) and space (place of collection).
ONLINE RESOURCES
NCBI (National Center for Biotechnology Information)
Site: http://www.ncbi.nlm.nih.gov/
Used: Daily
Official Description:
Established in 1988 as a national resource for molecular biology information, NCBI creates public databases, conducts research in computational biology, develops software tools for analyzing genome data, and disseminates biomedical information - all for the better understanding of molecular processes affecting human health and disease.
What I use it for:
PubMed serves as a primary reference search tool which is linked to the U.S. National Library of Medicine. NCBI houses several major databases, ranging from nucleotide and protein, to taxonomic and structure/function. NCBI also has databases dedicated to SNP (Single Nucleotide Polymorphim), EST (Expressed Sequence Tag), and GEO (Gene Expression Omnibus) analysis. Databases cover all forms of life, from eukaryotic (animal and plant) to prokaryotic (archaea and eubacterial). NCBI also serves the BLAST (Basic Local Alignment Search Tool) which is used to examine sequence similarity to other previously identified sequences (nucleotide or protein).
Ribosomal Database Project (Michigan State University, J.M. Tiedje)
Site: https://rdp.cme.msu.edu/
Used: Daily
Official Description:
The Ribosomal Database Project (RDP) provides ribosome related data and services to the scientific community, including online data analysis and aligned and annotated Bacterial small-subunit 16S rRNA sequences.
What I use it for:
Upon sequencing 16S clones, we use the RDP database to classify (Phylum/Class/Order/Family/Genus/Species) them for separation, for further phylogenetic analysis. The RDP also has sequence match functions which will identify closely related sequences which are useful when building phylogenetic trees (typically using MEGA).
Bellerophon
Site: http://foo.maths.uq.edu.au/~huber/bellerophon.pl
Used: Daily
Official Description:
Bellerophon is a program for detecting chimeric sequences in a multiple sequence dataset by comparative analysis. Bellerophon was specifically developed to detect 16S rRNA gene chimeras in PCR-clone libraries but can be applied to other gene datasets. A chimeric sequence, or chimera for short, is a sequence comprised of two or more phylogenetically distinct parent sequences. Chimeras are usually PCR artifacts thought to occur when a prematurely terminated amplicon reanneals to a foreign DNA strand and is copied to completion in the following PCR cycles. The point at which the chimeric sequence changes from one parent to the next is called the breakpoint or conversion point.
What I use it for:
Chimera detection in 16S sequences.
Here's my list:
IN HOUSE
Geneious ($249 subscription/year)
Used: Daily
Official description:
Geneious Pro is an integrated, cross-platform bioinformatics software suite for manipulating, finding, sharing, and exploring biological data such as DNA sequences or proteins, phylogenies, 3D structure information, publications, etc. It features sequence alignment and phylogenetic analysis, contig assembly, primer design and restriction analysis, access to NCBI and UniProt, BLAST, protein structure viewing, automated PubMed searching, and more. It even includes an API for creating your own plugins.
What I use it for:
Geneious is the workhorse application for DNA sequence analysis (chromatogram/sequence quality) and editing (vector and quality trimming) in my laboratory. Geneious is also used for contig assembly of genes/organisms and for alignment of 16S sequences for downstream phylogeny analysis (see programs MEGA, DnaSP, DAMBE). It can also be used to construct quick phylogenetic trees for routine examination. The subscription package allows me to receive regular updates. The only other comparable software application that I’ve found that works well on Windows XP is Sequencher (2007 quote for purchase was $2975. Major updates would require another purchase).
Artemis (freeware)
Used: Moderately (several times a month)
Official description:
Artemis is a free genome viewer and annotation tool that allows visualization of sequence features and the results of analyses within the context of the sequence, and its six-frame translation. Artemis is written in Java, and is available for UNIX, GNU/Linux, BSD, Macintosh and MS Windows systems. It can read complete EMBL and GENBANK database entries or sequence in FASTA or raw format. Extra sequence features can be in EMBL, GENBANK or GFF format.
What I use it for:
Artemis is a valuable tool for examining completed genomes. Search by gene/sequence/functional category for items of interest. GenBank houses over 630 completed microbial genomes (631 as of 02/08/08).
MEGA ver4.0 (freeware)
Used: Moderately
Official description:
MEGA is an integrated tool for conducting automatic and manual sequence alignment, inferring phylogenetic trees, mining web-based databases, estimating rates of molecular evolution, and testing evolutionary hypotheses.
What I use it for:
MEGA is the primary phylogenetic tree building program. It constructs publication quality phylogenetic trees. It is used for molecular evolution and population genetic analysis. In terms of alignment data, MEGA is an established format and most programs export/import alignments in MEGA format. Geneious exports alignment data in MEGA format, allowing these two programs to be used in conjunction. MEGA also has a sequence editor for quick/minor alignment editing.
DnaSP (freeware)
Used: Infrequently/Rarely
Official description:
DnaSP, DNA Sequence Polymorphism, is a software package for the analysis of nucleotide polymorphism from aligned DNA sequence data. DnaSP can estimate several measures of DNA sequence variation within and between populations (in noncoding, synonymous or nonsynonymous sites, or in various sorts of codon positions), as well as linkage disequilibrium, recombination, gene flow and gene conversion parameters. DnaSP can also carry out several tests of neutrality: Hudson, Kreitman and Aguadé, Tajima, McDonald and Kreitman, Fu and Li, and Fu tests. Additionally, DnaSP can estimate the confidence intervals of some test-statistics by the coalescent. The results of the analyses are displayed on tabular and graphic form.
What I use it for:
DnaSP is primarily used to determine genotype numbers. Genotypes are based on SNP information derived from DNA sequencing (typically MLST – multi-locus sequence typing) closely related strains/isolates.
DAMBE (freeware)
Used: Infrequently/Rarely
Official description:
Data analysis in molecular biology and evolution. t is an integrated software package for retrieving, organizing, manipulating, aligning, and analyzing molecular sequence data. Allele frequency data can also be used by DAMBE for calculating genetic distances or phylogenetic reconstruction.
What I use it for:
DAMBE does not see frequent usage in the lab, but it is sometimes useful for determining genotype numbers (it ignores gapped sequences for example) from complex sequences.
TotalLab 120 DM ($6,000)
Used: Moderately
Official description:
The TL120 version in the TotalLab range is an advanced image analysis solution which offers an extensive range of features for the in-depth analysis of 1D electrophoresis gels and performing band pattern matching studies. TL120 DM is the TL120 analysis software complete with the DM database component so you can archive all your analysed results and perform cross experiment investigations.
What I use it for:
This program is integral for analysis of our ribosomal intergenic spacer analysis (RISA) data which is collected on a LiCor DNA sequencer. This allows us to look at archaea, eubacterial and fungal population patterns in a sample and then compare that gel image to other images/samples to construct a phylogenetic relationship between them. The DM option allows us to store this information in a database for comparison of data between experiments. This will enhance our ability to compare samples across time (date of analysis & time of collection) and space (place of collection).
ONLINE RESOURCES
NCBI (National Center for Biotechnology Information)
Site: http://www.ncbi.nlm.nih.gov/
Used: Daily
Official Description:
Established in 1988 as a national resource for molecular biology information, NCBI creates public databases, conducts research in computational biology, develops software tools for analyzing genome data, and disseminates biomedical information - all for the better understanding of molecular processes affecting human health and disease.
What I use it for:
PubMed serves as a primary reference search tool which is linked to the U.S. National Library of Medicine. NCBI houses several major databases, ranging from nucleotide and protein, to taxonomic and structure/function. NCBI also has databases dedicated to SNP (Single Nucleotide Polymorphim), EST (Expressed Sequence Tag), and GEO (Gene Expression Omnibus) analysis. Databases cover all forms of life, from eukaryotic (animal and plant) to prokaryotic (archaea and eubacterial). NCBI also serves the BLAST (Basic Local Alignment Search Tool) which is used to examine sequence similarity to other previously identified sequences (nucleotide or protein).
Ribosomal Database Project (Michigan State University, J.M. Tiedje)
Site: https://rdp.cme.msu.edu/
Used: Daily
Official Description:
The Ribosomal Database Project (RDP) provides ribosome related data and services to the scientific community, including online data analysis and aligned and annotated Bacterial small-subunit 16S rRNA sequences.
What I use it for:
Upon sequencing 16S clones, we use the RDP database to classify (Phylum/Class/Order/Family/Genus/Species) them for separation, for further phylogenetic analysis. The RDP also has sequence match functions which will identify closely related sequences which are useful when building phylogenetic trees (typically using MEGA).
Bellerophon
Site: http://foo.maths.uq.edu.au/~huber/bellerophon.pl
Used: Daily
Official Description:
Bellerophon is a program for detecting chimeric sequences in a multiple sequence dataset by comparative analysis. Bellerophon was specifically developed to detect 16S rRNA gene chimeras in PCR-clone libraries but can be applied to other gene datasets. A chimeric sequence, or chimera for short, is a sequence comprised of two or more phylogenetically distinct parent sequences. Chimeras are usually PCR artifacts thought to occur when a prematurely terminated amplicon reanneals to a foreign DNA strand and is copied to completion in the following PCR cycles. The point at which the chimeric sequence changes from one parent to the next is called the breakpoint or conversion point.
What I use it for:
Chimera detection in 16S sequences.
Subscribe to:
Posts (Atom)