Monday, January 10, 2011

SMILES...

SMILES - A Simplified Chemical Language
SMILESTM
Simplified Molecular Input Line Entry System

SMILESTM as a simple yet comprehensive chemical language in which molecules and reactions can be specified using ASCII characters representing atom and bond symbols. SMILESTM contains the same information as is found in an extended connection table but with several advantages. A SMILESTM string is human understandable, very compact, and if canonicalized represents a unique string that can be used as a universal identifier for a specific chemical structure. In addition, a chemically correct and comprehensible depiction can be made from any SMILESTM string symbolizing either a molecule or reaction.

SMILESTM development was initiated by David Weininger in the late 1980s using the concept of a graph with nodes as atoms and edges as bonds to represent a molecule. Parentheses are used to indicate branching points and numeric labels designate ring connection points. The basic SMILESTM grammar also includes as well as isotopic information, configuration about double bonds, and chirality leading to what is known as isomeric SMILESTM.

Some simple SMILESTM examples:
Ethanol CCO
Acetic acid CC(=O)O
Cyclohexane C1CCCCC1
Pyridine c1cnccc1
Trans-2-butene C/C=C/C
L-alanine N[C@@H](C)C(=O)O
Sodium chloride [Na+].[Cl-]
Displacement reaction     C=CCBr>>C=CCI
Since its inception, SMILESTM has been modified and expanded by Daylight to include not only new features but two additional chemical languages: SMARTS®, an expansion of SMILESTM allowing specification of molecular patterns and properties for substructure searching with varying levels of specificity, and SMIRKS®, a restricted version of reaction SMARTS® involving changes in atom-bond patterns that define generic reactions. 
 
SMILES (Simplified Molecular Input Line Entry System) is a line notation (a typographical method using printable characters) for entering and representing molecules and reactions. Some examples are:

SMILESNameSMILESName
CC ethane [OH3+] hydronium ion
O=C=O carbon dioxide [2H]O[2H] deuterium oxide
C#N hydrogen cyanide [235U] uranium-235
CCN(CC)CC triethylamine F/C=C/F E-difluoroethene
CC(=O)O acetic acid F/C=C\F Z-difluoroethene
C1CCCCC1 cyclohexane N[C@@H](C)C(=O)O L-alanine
c1ccccc1 benzene N[C@H](C)C(=O)O D-alanine


Reaction SMILES Name
[I-].[Na+].C=CCBr>>[Na+].[Br-].C=CCI displacement reaction
(C(=O)O).(OCC)>>(C(=O)OCC).(O) intermolecular esterification


SMILES contains the same information as might be found in an extended connection table. The primary reason SMILES is more useful than a connection table is that it is a linguistic construct, rather than a computer data structure. SMILES is a true language, albeit with a simple vocabulary (atom and bond symbols) and only a few grammar rules. SMILES representations of structure can in turn be used as "words" in the vocabulary of other languages designed for storage of chemical information (information about chemicals) and chemical intelligence (information about chemistry).

Part of the power of SMILES is that unique SMILES exist. With standard SMILES, the name of a molecule is synonymous with its structure; with unique SMILES, the name is universal. Anyone in the world who uses unique SMILES to name a molecule will choose the exact same name.

One other important property of SMILES is that it is quite compact compared to most other methods of representing structure. A typical SMILES will take 50% to 70% less space than an equivalent connection table, even binary connection tables. For example, a database of 23,137 structures, with an average of 20 atoms per structure, uses only 1.6 bytes per atom when represented with SMILES. In addition, ordinary compression of SMILES is extremely effective. The same database cited above was reduced to 27% of its original size by Ziv-Lempel compression (i.e. 0.42 bytes per atom).

These properties open many doors to the chemical information programmer. Examples of uses for SMILES are:

  • Keys for database access
  • Mechanism for researchers to exchange chemical information
  • Entry system for chemical data
  • Part of languages for artificial intelligence or expert systems in chemistry

The rest of this chapter is a concise exposition of the SMILES encoding rules. For further information, the reader is referred to "SMILES 1. Introduction and Encoding Rules", Weininger, D., J.Chem. Inf. Comput. Sci. 1988, 28,31.

Presented to you SMILES....



Sunday, January 2, 2011

Protein Data Bank....PDB

Protein Data Bank???

The Protein Data Bank (PDB) is a repository for the 3-D structural data of large biological molecules, such as proteins and nucleic acids. (See also crystallographic database). The data, typically obtained by X-ray crystallography or NMR spectroscopy and submitted by biologists and biochemists from around the world, are freely accessible on the Internet via the websites of its member organisations (PDBe, PDBj, and RCSB). The PDB is overseen by an organization called the Worldwide Protein Data Bank, wwPDB.

The PDB is a key resource in areas of structural biology, such as structural genomics. Most major scientific journals, and some funding agencies, such as the NIH in the USA, now require scientists to submit their structure data to the PDB. If the contents of the PDB are thought of as primary data, then there are hundreds of derived (i.e., secondary) databases that categorize the data differently. For example, both SCOP and CATH categorize structures according to type of structure and assumed evolutionary relations; GO categorize structures based on genes.

Viewing the data

The structure files may be viewed using one of several open source computer programs. Some other free, but not open source programs include VMD, MDL Chime, Swiss-PDB Viewer, StarBiochem (a Java-based interactive molecular viewer with integrated search of protein databank), Sirius, and VisProt3DS (a tool for Protein Visualization in 3D stereoscopic view in anaglyth and other modes). The RCSB PDB website contains an extensive list of both free and commercial molecule visualization programs and web browser plugins.

These samples that we can get from PDB :

HtrA

3LT3
Crystal structure of Rv3671c from M. tuberculosis H37Rv, Ser343Ala mutant, inactive form
Authors:
Biswas, T., Small, J., Vandal, O., Ehrt, S., Tsodikov, O.V.
Release Date: 2010-11-03
Classification: Hydrolase
Experiment: X-RAY DIFFRACTION with resolution of 2.10 Å
Compound: 1 Polymer
Molecule: POSSIBLE MEMBRANE-ASSOCIATED SERINE PROTEASE
Polymer: 1  
Type: polypeptide(L) Length: 217
Chains: A, B
EC#: 3.4.21.-
Fragment: Rv3671c (179-397)
Mutation: S343A
Citation: Structural insight into serine protease Rv3671c that Protects M. tuberculosis from oxidative and acidic stress.
(2010) Structure 18: 1353-1363

PubMed Abstract:
Rv3671c, a putative serine protease, is crucial for persistence of Mycobacterium tuberculosis in the hostile environment of the phagosome. We show that Rv3671c is required for M. tuberculosis resistance to oxidative stress in addition to its role in protection from acidification. Structural and biochemical analyses demonstrate that the periplasmic domain of Rv3671c is a functional serine protease of the chymotrypsin family and, remarkably, that its activity increases on oxidation. High-resolution crystal structures of this protease in an active strained state and in an inactive relaxed state reveal that a solvent-exposed disulfide bond controls the protease activity by constraining two distant regions of Rv3671c and stabilizing it in the catalytically active conformation. In vitro biochemical studies confirm that activation of the protease in an oxidative environment is dependent on this reversible disulfide bond. These results suggest that the disulfide bond modulates activity of Rv3671c depending on the oxidative environment in vivo.

Citation Authors:
Biswas, T., Small, J., Vandal, O., Odaira, T., Deng, H., Ehrt, S., Tsodikov, O.V.


LonA

3M65
Crystal structure of Bacillus subtilis Lon N-terminal domain
Authors:
Duman, R.E., Lowe, J.Y.
Release Date: 2010-06-30
Classification: Hydrolase
Experiment: X-RAY DIFFRACTION with resolution of 2.60 Å
Compound: 1 Polymer
Molecule: ATP-dependent protease La 1
Polymer: 1 Type: polypeptide(L)  
Length: 209
Chains: A, B
EC#: 3.4.21.53
Fragment: BsLon, N-terminal domain
Citation: Crystal Structures of Bacillus subtilis Lon Protease.
(2010) J.Mol.Biol. 401: 653-670

PubMed Abstract:
Lon ATP-dependent proteases are key components of the protein quality control systems of bacterial cells and eukaryotic organelles. Eubacterial Lon proteases contain an N-terminal domain, an ATPase domain, and a protease domain, all in one polypeptide chain. The N-terminal domain is thought to be involved in substrate recognition, the ATPase domain in substrate unfolding and translocation into the protease chamber, and the protease domain in the hydrolysis of polypeptides into small peptide fragments. Like other AAA+ ATPases and self-compartmentalising proteases, Lon functions as an oligomeric complex, although the subunit stoichiometry is currently unclear. Here, we present crystal structures of truncated versions of Lon protease from Bacillus subtilis (BsLon), which reveal previously unknown architectural features of Lon complexes. Our analytical ultracentrifugation and electron microscopy show different oligomerisation of Lon proteases from two different bacterial species, Aquifex aeolicus and B. subtilis. The structure of BsLon-AP shows a hexameric complex consisting of a small part of the N-terminal domain, the ATPase, and protease domains. The structure shows the approximate arrangement of the three functional domains of Lon. It also reveals a resemblance between the architecture of Lon proteases and the bacterial proteasome-like protease HslUV. Our second structure, BsLon-N, represents the first 209 amino acids of the N-terminal domain of BsLon and consists of a globular domain, similar in structure to the E. coli Lon N-terminal domain, and an additional four-helix bundle, which is part of a predicted coiled-coil region. An unexpected dimeric interaction between BsLon-N monomers reveals the possibility that Lon complexes may be stabilised by coiled-coil interactions between neighbouring N-terminal domains. Together, BsLon-N and BsLon-AP are 36 amino acids short of offering a complete picture of a full-length Lon protease.

Citation Authors:
Duman, R.E., Lowe, J.

ClpP

3MT6
Structure of ClpP from Escherichia coli in complex with ADEP1
Authors:
Chung, Y.S.
Release Date: 2010-11-03  
Classification: Hydrolase/antibiotic
Experiment: X-RAY DIFFRACTION with resolution of 1.90 Å
Compound: 2 Polymers
Molecule: ATP-dependent Clp protease proteolytic subunit
Polymer: 1 Type: polypeptide(L)  
Length: 207
Chains: A, B, C, D, E, F, G, H, I, J, K, L, M, N, O, P, Q, R, S, T, U, V, W, X, Y, Z, a, b
EC#: 3.4.21.92
Molecule: ACYLDEPSIPEPTIDE 1
Polymer: 2     
Type: polypeptide(L) Length: 7
Chains: 1, 2, 3, 4, c, d, e, f, g, h, i, j, k, l, m, n, o, p, q, r, s, t, u, v, w, x, y, z

1 Ligand
Image Identifier Name Formula
MPD   (4S)-2-METHYL-2,4-PENTANEDIOL C6 H14 O2

PubMed Abstract:
In ClpXP and ClpAP complexes, ClpA and ClpX use the energy of ATP hydrolysis to unfold proteins and translocate them into the self-compartmentalized ClpP protease. ClpP requires the ATPases to degrade folded or unfolded substrates, but binding of acyldepsipeptide antibiotics (ADEPs) to ClpP bypasses this requirement with unfolded proteins. We present the crystal structure of Escherichia coli ClpP bound to ADEP1 and report the structural changes underlying ClpP activation. ADEP1 binds in the hydrophobic groove that serves as the primary docking site for ClpP ATPases. Binding of ADEP1 locks the N-terminal loops of ClpP in a ?-hairpin conformation, generating a stable pore through which extended polypeptides can be threaded. This structure serves as a model for ClpP in the holoenzyme ClpAP and ClpXP complexes and provides critical information to further develop this class of antibiotics.

Citation Authors:
Li, D.H., Chung, Y.S., Gloyd, M., Joseph, E., Ghirlando, R., Wright, G.D., Cheng, Y.Q., Maurizi, M.R., Guarne, A., Ortega, J.