Bioinformatics Tools

Pages

Showing posts with label Bioinformatics. Show all posts
Showing posts with label Bioinformatics. Show all posts

Thursday, July 16, 2015

Woods: A fast and accurate functional annotator and classifier of genomic and metagenomic sequences

Another recent publication from our lab.

Citation: 

Sharma, A. K., Gupta, A., Kumar, S., Dhakan, D. B., & Sharma, V. K. (2015). Woods: A fast and accurate functional annotator and classifier of genomic and metagenomic sequences. Genomics.

Functional annotation of the gigantic metagenomic data is one of the major time-consuming and computationally demanding tasks, which is currently a bottleneck for the efficient analysis. The commonly used homology-based methods to functionally annotate and classify proteins are extremely slow. 

Therefore, to achieve faster and accurate functional annotation, we have developed an orthology-based functional classifier 'Woods' by using a combination of machine learning and similarity-based approaches. Woods displayed a precision of 98.79% on independent genomic dataset, 96.66% on simulated metagenomic dataset and >97% on two real metagenomic datasets. In addition, it performed >87 times faster than BLAST on the two real metagenomic datasets. Woods can be used as a highly efficient and accurate classifier with high-throughput capability which facilitates its usability on large metagenomic datasets.

The Woods web server is freely accessible at http://metagenomics.iiserb.ac.in/woods/index.php and http://metabiosys.iiserb.ac.in/woods/index.php. The standalone version of Woods can be downloaded from the above web servers and usage instructions are provided in Text S1 and also in the Tutorial section of the web server.

Thursday, February 27, 2014

Short Simple Tips and Techniques for use in Bioinformatics

Here I am compiling some of the short simple tips and tricks for use to ease out some of Bioinformatics work, I would also be thankful if you can contribute to it, if you can come across some. Write them as comment and I will update them.

 #####################################

Running single or multiple commands on all files or selected files in a directory.


For running a command on number of files in a directory, use this:

CODE:

for i in *; do command; done;

Wherein * will select all the files to act upon, you can use *.fasta to work only with the fasta files, etc.
command is the linux command you wish to act on all the selected files, where in each selected file is represented by or recalled by "$i" and can also be used to generate output with the filename added to it.

Example:

for i in *.fasta; do grep ">gi" "$i" > "$i".output; done;

Similarly to copy a file from anywhere (previous directory in this example) in multiple directories

for i in *; do cp ../program.pl "$i" ; done;

##################################################

For working on multiples files, like removing, copying etc. on every directory inside a directory or multiple directories do the following: 

To remove all the fasta files from the folder or sub folders

find /pathtofolder/ -type f -iname "*.fasta" -exec rm {} \;

This will find all the files
find /pathtofolder/ -type f -iname "*.fasta"

and this will remove these files
 -exec rm {} \;

##################################################

To remove files containing a particular text in them from various folders inside the directory: 

For example you have downloaded all genomes *.fna from NCBI Genomes for Bacterial Genomes and you just wish to use *.fna from the bacterial genomes but not from the plasmids which you wish to remove.


grep -lrIZ plasmid . | xargs -0 rm -f --

This will remove all files containing plasmid in it.

##################################################

To grep the count (say fasta) in various files inside a directory: 

grep -c ">" */*/*.fasta > reads_count

This will count the number of reads from each file in those directories along with the file name and that can be redirected to a file.

NOTE: Change */*/*.fasta depending upon the number of directories inside i.e. */*.fasta will work if the files are in one directory below the present directory.

##################################################

 Calculating Mean Length of FASTA Sequences

 Here's a shot simple code for calculating the mean length of a FASTA sequence.

Code:


awk '{/>/&&++a||b+=length()}END{print b/a}'  sequence.fasta > average

Alternate faster methods are available, read here BIOSTARS

##################################################

Convert FASTQ to FASTA:

Code:

sed -n ' 1~4s/^@/>/p; 2~4p ' file.fastq > file.fasta


sed 's/\ /\_/g' seqfile > seqfile_ed (to make unique ids in flash assembled sequences)


##################################################

Schedule one job after another

Run one script after another in such a way that second script starts after finishing first one. Without using Pipe | or ampercent && i.e. the first process is already running and you want second one to start after the first one finishes. And this can be done in different folder in case the output of second script will affect the output of first script. So run this on any folder you wish to.

Code:

while ps -p $PID; do sleep 1; done; script2

Explanation and Example:

while ps -p 4437; do sleep 1; done; echo "It Works"

Where $PID is the process id of the already running job (add PID number)

script2 is your script you wish to run after first script ends

sleep 1 is sleep for one second (SUFFIX may be ‘s’ for seconds (the default), ‘m’ for minutes, ‘h’ for hours or ‘d’ for days, read man sleep)

##################################################

Sample command for formatdb and blastall

FORMAT_DB:
nohup time /home/sanjiv/Softwares/blast-2.2.26/bin/formatdb -p T -i input.database &

BLAST_ALL:

Protein BLAST (blastp)

nohup /home/sanjiv/Softwares/blast-2.2.22/bin/blastall -p blastp -d location_of_database.fa -i location_of_input_file.fasta -o  location_of_output_file.blastall.out -a 30 -e 0.000001 &

Nucleotide BLAST (blastx)

nohup /home/sanjiv/Softwares/blast-2.2.22/bin/blastall -p blastx -d location_of_database.fa -i location_of_input_file.fasta -o  location_of_output_file.blastall.out -a 30 -e 0.000001 &

##################################################

Fetch FASTA sequences using identifier from a file.

CODE:

grep -A1 -f list mainfile > fastaofidsfromlist

Explanation:

To fetch selected fasta sequence from huge file. Keep your identifiers or fasta headers in 'list' (i.e. like Rv ids). 'mainfile' is the original file from where you have to fetch the fasta. 'fastaofidsfromlist' is the your file with fasta sequences of the ids you gave as list. Issue the following command.

CAUTION: For this to work properly, your mainfile should have sequences in one line only, without line breaks. Check output file before use.

##################################################

Convert multi-line fasta to single line fasta

awk 'BEGIN{RS=">"}NR>1{sub("\n","\t"); gsub("\n",""); print RS$0}'  file

 ##################################################

Using BLAST+

Make blast database

makeblastdb -in proteases.faa -dbtype 'prot' -out CJProteases

BLAST

blastp -query proteins.fasta -db DATABASE -out OUTPUT_File -evalue 1e-30 -outfmt 7

Saturday, November 5, 2011

In-silico characterization of proteins

BLAST : In bioinformatics, Basic Local Alignment Search Tool, or BLAST, is an algorithm for comparing primary biological sequence information, such as the amino-acid sequences of different proteins or the nucleotides of DNA sequences. A BLAST search enables a researcher to compare a query sequence with a library or database of sequences, and identify library sequences that resemble the query sequence above a certain threshold. Different types of BLASTs are available according to the query sequences. For example, following the discovery of a previously unknown gene in the mouse, a scientist will typically perform a BLAST search of the human genome to see if humans carry a similar gene; BLAST will identify sequences in the human genome that resemble the mouse gene based on similarity of sequence. The BLAST program was designed by Eugene Myers, Stephen Altschul, Warren Gish, David J. Lipman, and Webb Miller at the NIH and was published in the Journal of Molecular Biology in 1990

CDD search: Conserved Domain Database (CDD) is a protein annotation resource that consists of a collection of well-annotated multiple sequence alignment models for ancient domains and full-length proteins. These are available as position-specific score matrices (PSSMs) for fast identification of conserved domains in protein sequences via RPS-BLAST. CDD content includes NCBI-curated domains, which use 3D-structure information to explicitly to define domain boundaries and provide insights into sequence/structure/function relationships, as well as domain models imported from a number of external source databases (Pfam, SMART, COG, PRK, TIGRFAM).

PFAM: The Pfam database is a large collection of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Proteins are generally composed of one or more functional regions, commonly termed domains. Different combinations of domains give rise to the diverse range of proteins found in nature. The identification of domains that occur within proteins can therefore provide insights into their function. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families. Although these Pfam-A entries cover a large proportion of the sequences in the underlying sequence database, in order to give a more comprehensive coverage of known proteins we also generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans. A clan is a collection of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM.

TMHMM: A variety of tools are available to predict the topology of transmembrane proteins. To date no independent evaluation of the performance of these tools has been published. A better understanding of the strengths and weaknesses of the different tools would guide both the biologist and the bioinformatician to make better predictions of membrane protein topology.

SignalP: SignalP 4.0 server predicts the presence and location of signal peptide cleavage sites in amino acid sequences from different organisms: Gram-positive prokaryotes, Gram-negative prokaryotes, and eukaryotes. The method incorporates a prediction of cleavage sites and a signal peptide/non-signal peptide prediction based on a combination of several artificial neural networks. 

STRING: STRING is a database of known and predicted protein interactions. The interactions include direct (physical) and indirect (functional) associations; they are derived from four sources i.e. Genomic context, high throughput experiments, coexpression, previous knowledge. STRING quantitatively integrates interaction data from these sources for a large number of organisms, and transfers information between these organisms where applicable. The database currently covers 5'214'234 proteins from 1133 organisms.

PROTPARAM: ProtParam (References / Documentation) is a tool which allows the computation of various physical and chemical parameters for a given protein stored in Swiss-Prot or TrEMBL or for a user entered sequence. The computed parameters include the molecular weight, theoretical pI, amino acid composition, atomic composition, extinction coefficient, estimated half-life, instability index, aliphatic index and grand average of hydropathicity (GRAVY)

PROSITE: Search your query sequence for protein motifs, rapidly compare your query protein sequence against all patterns stored in the PROSITE pattern database and determine what the function of an uncharacterised protein is. This tool requires a protein sequence as input, but DNA/RNA may be translated into a protein sequence using transeq and then queried.

InterPro: InterPro is an integrated database of predictive protein "signatures" used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures.

GlobPlot Webservice:

Prediction of disorder:

  • DisEMBL - DisEMBL is our neural network based predictor.
  • DISOPRED - Predictor from David Jones' lab.

Function prediction in non-globular protein space:

  • ELM - The Eukaryotic Linear Motif Resource.
  • NetworKIN - Systematic Discovery of In Vivo Phosphorylation Networks.

Thesis on disorder and linear motifs

Function prediction in globular protein space:

  • SMART - SMART/Pfam domains

Domain boundaries:

  • DomCut - A domain boundary detector
  • DomPred - Domain predictor from David Jones' lab.

Synthetic Biology

Synthetic Biology Project @ SLRI - Applying GlobPlot.


Resources

Subcellular localization predictors:
Subcellular localization databases: