Bioinformatics Tools

Pages

Wednesday, November 23, 2011

Folder list to text file, text file to folders

How to make folder with name from test file ?

You could do this:
1. Make sure all your entries are in column A of your spreadsheet.
2. Edit/copy column A
3. Click Start / Run / notepad c:\folders.txt {OK}
4. Click Edit / paste. You now have a text file with all the folder names
inside.
5. Click Start / run / cmd {OK}
6. Type this test command:
for /F "tokens=*" %* in (c:\folders.txt) do @echo md "D:\My Folders\%*"
{Enter}

If you're happy with the result, make it happen by typing this command:
for /F "tokens=*" %* in (c:\folders.txt) do @md "D:\My Folders\%*"
{Enter}



How do I print a listing of files in a directory?

    Get to the MS-DOS prompt or the Windows command line.
    Navigate to the directory you wish to print the contents of. If you're new to the command line, familiarize yourself with the cd command and the dir command.
    Once in the directory you wish to print the contents of, type one of the below commands.

    dir > print.txt

    The above command will take the list of all the files and all of the information about the files, including size, modified date, etc., and send that output to the print.txt file in the current directory.

    dir /b > print.txt

    This command would print only the file names and not the file information of the files in the current directory.

    dir /s /b > print.txt

    This command would print only the file names of the files in the current directory and any other files in the directories in the current directory.

    After doing any of the above steps the print.txt file is created. Open this file in any text editor (e.g. Notepad) and print the file. You can also do this from the command prompt by typing notepad print.txt.

Saturday, November 5, 2011

In-silico characterization of proteins

BLAST : In bioinformatics, Basic Local Alignment Search Tool, or BLAST, is an algorithm for comparing primary biological sequence information, such as the amino-acid sequences of different proteins or the nucleotides of DNA sequences. A BLAST search enables a researcher to compare a query sequence with a library or database of sequences, and identify library sequences that resemble the query sequence above a certain threshold. Different types of BLASTs are available according to the query sequences. For example, following the discovery of a previously unknown gene in the mouse, a scientist will typically perform a BLAST search of the human genome to see if humans carry a similar gene; BLAST will identify sequences in the human genome that resemble the mouse gene based on similarity of sequence. The BLAST program was designed by Eugene Myers, Stephen Altschul, Warren Gish, David J. Lipman, and Webb Miller at the NIH and was published in the Journal of Molecular Biology in 1990

CDD search: Conserved Domain Database (CDD) is a protein annotation resource that consists of a collection of well-annotated multiple sequence alignment models for ancient domains and full-length proteins. These are available as position-specific score matrices (PSSMs) for fast identification of conserved domains in protein sequences via RPS-BLAST. CDD content includes NCBI-curated domains, which use 3D-structure information to explicitly to define domain boundaries and provide insights into sequence/structure/function relationships, as well as domain models imported from a number of external source databases (Pfam, SMART, COG, PRK, TIGRFAM).

PFAM: The Pfam database is a large collection of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Proteins are generally composed of one or more functional regions, commonly termed domains. Different combinations of domains give rise to the diverse range of proteins found in nature. The identification of domains that occur within proteins can therefore provide insights into their function. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families. Although these Pfam-A entries cover a large proportion of the sequences in the underlying sequence database, in order to give a more comprehensive coverage of known proteins we also generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans. A clan is a collection of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM.

TMHMM: A variety of tools are available to predict the topology of transmembrane proteins. To date no independent evaluation of the performance of these tools has been published. A better understanding of the strengths and weaknesses of the different tools would guide both the biologist and the bioinformatician to make better predictions of membrane protein topology.

SignalP: SignalP 4.0 server predicts the presence and location of signal peptide cleavage sites in amino acid sequences from different organisms: Gram-positive prokaryotes, Gram-negative prokaryotes, and eukaryotes. The method incorporates a prediction of cleavage sites and a signal peptide/non-signal peptide prediction based on a combination of several artificial neural networks. 

STRING: STRING is a database of known and predicted protein interactions. The interactions include direct (physical) and indirect (functional) associations; they are derived from four sources i.e. Genomic context, high throughput experiments, coexpression, previous knowledge. STRING quantitatively integrates interaction data from these sources for a large number of organisms, and transfers information between these organisms where applicable. The database currently covers 5'214'234 proteins from 1133 organisms.

PROTPARAM: ProtParam (References / Documentation) is a tool which allows the computation of various physical and chemical parameters for a given protein stored in Swiss-Prot or TrEMBL or for a user entered sequence. The computed parameters include the molecular weight, theoretical pI, amino acid composition, atomic composition, extinction coefficient, estimated half-life, instability index, aliphatic index and grand average of hydropathicity (GRAVY)

PROSITE: Search your query sequence for protein motifs, rapidly compare your query protein sequence against all patterns stored in the PROSITE pattern database and determine what the function of an uncharacterised protein is. This tool requires a protein sequence as input, but DNA/RNA may be translated into a protein sequence using transeq and then queried.

InterPro: InterPro is an integrated database of predictive protein "signatures" used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures.

GlobPlot Webservice:

Prediction of disorder:

  • DisEMBL - DisEMBL is our neural network based predictor.
  • DISOPRED - Predictor from David Jones' lab.

Function prediction in non-globular protein space:

  • ELM - The Eukaryotic Linear Motif Resource.
  • NetworKIN - Systematic Discovery of In Vivo Phosphorylation Networks.

Thesis on disorder and linear motifs

Function prediction in globular protein space:

  • SMART - SMART/Pfam domains

Domain boundaries:

  • DomCut - A domain boundary detector
  • DomPred - Domain predictor from David Jones' lab.

Synthetic Biology

Synthetic Biology Project @ SLRI - Applying GlobPlot.


Resources

Subcellular localization predictors:
Subcellular localization databases:

Wednesday, February 24, 2010

Saturday, May 16, 2009

Wolfram|Alpha: Future of search and analysis

This is something very interesting I came across recently. I have just seen the DEMO which is so mind blowing, actual site is located here Wolfram|Alpha. God!!! we are a growing species, I just have to say that it is just mind blowing....marvelous..yes we human ...mind is the speck over the hell to achieve unimaginable through the labyrinthine network of our neuronal circuit ..its evitable the more guttural pouch will emerge like this..which will be more useful in days to come...This is the future of the way we look at the internet today, the search we do over internet, The search results and their analysis, which otherwise takes us to 100s of pages with analysis for the required details over each page. Anyways, this Wolfram|Alpha looks so good and I am wondering that it came bit late. Very amazing piece of work, I am sure this is going to be very very useful to almost every internet user. I am also hopeful that as in comparison to the Google search results ,this gives more analyzed results. And no I am not comparing these two. Lots of mathematical calculation, lots of general work analysis. I am sure Wolfram|Alpha people are and will be having tough time maintaining such a brilliant technology.

Tuesday, March 31, 2009

miRex: A web based resource for miRNA expression profiles

Background: A few hundred miRNAs carry the potential to regulate thousands of target genes in eukaryotes. The expression profiles of miRNAs convey important information regarding tissue specific gene expression and can be used as a biomarker for disease progression and cancer classification among other rational interpretations pertaining to miRNA-gene interactions. There are several individual reports of miRNA expression profiling; however there is a lack of server that can render cross-comparison of all these datasets.


Description: We have developed miRex, a database and analysis tool for comparing miRNA expression profiles generated by high-throughput methods. Currently data from public repositories have been pre-normalized and provided with visual representation to aid comparison between experiments. miRNA ID converter: a tool for mapping miRNA IDs from one system of nomenclature to another has also been included.


Data: Currently, 614 experiments spanning 25 datasets deposited in Gene Expression Omnibus (GEO),the public repository for high-throughput gene expression data hosted by NCBI and 1132 experiments from 18 datasets from ArrayExpress, another resource for expression data, is available through miRex. Besides the microarray based data, there is a set of 40 experiments carried out by real time PCR.


URL: miRex is available at http://miracle.igib.res.in/mirex/


Wednesday, September 24, 2008

Reverse Complement

Reverse Complement

Reverse Complement is commonly used in Bioinformatics for various purposes. Here is the tool that does the job without much effort, there are simple Perl programs that could be run locally for the purpose. This tool is provided by GENE INFINITY, this can also do reverse and complementary separately. Hope this helps, the tool is located here, Reverse Complement

Protein Blast against another set of proteins

Protein Blast against another set of proteins

This tool is provided by NCBI/ BLAST/ blastp suite: BLASTP programs search protein databases using a protein query.This gives BLAST of a query protein against a set of other proteins. I found it useful when you don't wish to BLAST your query against whole protein database, instead a set of proteins given by the user. This tool is located here, Protein Blast against another set of proteins

PeptideCutter

PeptideCutterhttp://expasy.org/tools/peptidecutter/

This tool is provided by ExPASy. This predicts potential cleavage sites cleaved by proteases or chemicals in a given protein sequence. 

PeptideCutter returns the query sequence with the possible cleavage sites mapped on it and /or a table of cleavage site positions. Single or multiple enzymes can be selected for the purpose. PeptideCutter

Predicting Antigenic Peptides

Predicting Antigenic Peptides

This is a program that predicts those segments from within a protein sequence that are likely to be antigenic by eliciting an antibody response. The method used here is the method of Kolaskar and Tongaonkar (1990). 

Predictions are based on a table that reflects the occurrence of amino acid residues in experimentally known segmental epitopes. Segments are only reported if the have a minimum size of 8 residues. The reported accuracy of method is about 75%. 

The program is located here Predicting Antigenic Peptides

Friday, February 1, 2008

Sequence analyzer

Sequence Massagerhttp://www.attotron.com/cybertory/analysis/seqMassager.htm

Nucleic Acid Sequence Massager is a very easy to use tool for convention of DNA to RNA, RNA to DNA, Upper Case to Lower Case and vice verse, Removal of FASTA format, Removal of HTML tags, Removal of number, White spaces, line breaks.

I find this tool very handy.

Saturday, September 15, 2007

Predicting Subcellular Localization of Proteins

It is interesting to study the localization of proteins in subcellular due to several reasons. Here is a collection of the online available softwares that help in predicting subcellular localization of the proteins. Prediction is done with the help of programs which are trained for this purpose, this greatly helps in selection procedure, to select for a protein to work upon. Though there are more I have enlisted some commonly used.
CELLO : CELLO is a multi-class SVM classification system. CELLO uses 4 types of sequence coding schemes: the amino acid composition, the di-peptide composition, the partitioned amino acid composition and the sequence composition based on the physico-chemical properties of amino acids. We combine votes from these classifiers and use the jury votes to determine the final assignment. Yu CS, Lin CJ, Hwang JK: Predicting subcellular localization of proteins for Gram-negative bacteria by support vector machines based on n-peptide compositions. Protein Science 2004, 13:1402-1406.
PSORTb: Based on a study last performed in 2010, PSORTb v3.0.2 is the most precise bacterial localization prediction tool available. PSORTb v3.0.2 has a number of improvements over PSORTb v2.0.4. Version 2 of PSORTb is maintained here. You can currently submit one or more Gram-positive or Gram-negative bacterial sequences or archaeal sequences in FASTA format. Copy and paste your FASTA-formatted sequences into the textbox below or select a file containing your sequences to upload from your computer.


TMHMM Server: This server is for prediction of transmembrane helices in proteins. You can submit many proteins at once in one fasta file. Please limit each submission to at most 4000 proteins. Please tick the 'One line per protein' option. Please leave time between each large submission.S. Moller, M.D.R. Croning, R. Apweiler. Evaluation of methods for the prediction of membrane spanning regions. Bioinformatics, 17(7):646-653, July 2001.

SignalP 3.0 Server: SignalP 3.0 server predicts the presence and location of signal peptide cleavage sites in amino acid sequences from different organisms: Gram-positive prokaryotes, Gram-negative prokaryotes, and eukaryotes. The method incorporates a prediction of cleavage sites and a signal peptide/non-signal peptide prediction based on a combination of several artificial neural networks and hidden Markov models. Locating proteins in the cell using TargetP, SignalP, and related tools Olof Emanuelsson, Søren Brunak, Gunnar von Heijne, Henrik Nielsen Nature Protocols 2, 953-971 (2007).


LOCtree: LOCtree can predict the subcellular localization and DNA-binding propensity of non-membrane proteins in non-plant and plant eukaryotes as well as prokaryotes. LOCtree classifies eukaryotic animal proteins into one of five subcellular classes, while plant proteins are classified into one of six classes and prokaryotic proteins are classified into one of three classes . The novel feature of using a hierarchical architecture is the ability to make intermediate localization class predictions at much higher accuracy's. Another source of improvement is the use of 'noisy' training data. 'Noisy' predictions from LOCKey (SWISS-PROT keyword based annotations) and LOCHom (annotations using sequence homology) are used to train the hierarchical SVMs.


PredictProtein: PredictProtein integrates feature prediction for secondary structure, solvent accessibility, transmembrane helices, globular regions, coiled-coil regions ,structural switch regions, B-values, disorder regions, intra-residue contacts, protein-protein and protein-DNA binding sites, sub-cellular localization, domain boundaries, beta-barrels, cysteine bonds, metal binding sites and disulphide bridges.