Mostrando entradas con la etiqueta aminoácidos. Mostrar todas las entradas
Mostrando entradas con la etiqueta aminoácidos. Mostrar todas las entradas

lunes, 13 de enero de 2014

Computer science: The learning machines

Using massive amounts of data to recognize photos and speech, deep-learning computers are taking a big step towards true artificial intelligence.

Article tools
PDF
Rights & Permissions

BRUCE ROLFF/SHUTTERSTOCK

Three years ago, researchers at the secretive Google X lab in Mountain View, California, extracted some 10 million still images from YouTube videos and fed them into Google Brain — a network of 1,000 computers programmed to soak up the world much as a human toddler does. After three days looking for recurring patterns, Google Brain decided, all on its own, that there were certain repeating categories it could identify: human faces, human bodies and … cats1.

Google Brain's discovery that the Internet is full of cat videos provoked a flurry of jokes from journalists. But it was also a landmark in the resurgence of deep learning: a three-decade-old technique in which massive amounts of data and processing power help computers to crack messy problems that humans solve almost intuitively, from recognizing faces to understanding language.

Deep learning itself is a revival of an even older idea for computing: neural networks. These systems, loosely inspired by the densely interconnected neurons of the brain, mimic human learning by changing the strength of simulated neural connections on the basis of experience. Google Brain, with about 1 million simulated neurons and 1 billion simulated connections, was ten times larger than any deep neural network before it. Project founder Andrew Ng, now director of the Artificial Intelligence Laboratory at Stanford University in California, has gone on to make deep-learning systems ten times larger again.

Such advances make for exciting times in artificial intelligence (AI) — the often-frustrating attempt to get computers to think like humans. In the past few years, companies such as Google, Apple and IBM have been aggressively snapping up start-up companies and researchers with deep-learning expertise. For everyday consumers, the results include software better able to sort through photos, understand spoken commands and translate text from foreign languages. For scientists and industry, deep-learning computers can search for potential drug candidates, map real neural networks in the brain or predict the functions of proteins.

AI has gone from failure to failure, with bits of progress. This could be another leapfrog,” says Yann LeCun, director of the Center for Data Science at New York University and a deep-learning pioneer.

Over the next few years we'll see a feeding frenzy. Lots of people will jump on the deep-learning bandwagon,” agrees Jitendra Malik, who studies computer image recognition at the University of California, Berkeley. But in the long term, deep learning may not win the day; some researchers are pursuing other techniques that show promise. “I'm agnostic,” says Malik. “Over time people will decide what works best in different domains.

Inspired by the brain

Back in the 1950s, when computers were new, the first generation of AI researchers eagerly predicted that fully fledged AI was right around the corner. But that optimism faded as researchers began to grasp the vast complexity of real-world knowledge — particularly when it came to perceptual problems such as what makes a face a human face, rather than a mask or a monkey face. Hundreds of researchers and graduate students spent decades hand-coding rules about all the different features that computers needed to identify objects. “Coming up with features is difficult, time consuming and requires expert knowledge,” says Ng. “You have to ask if there's a better way.

IMAGES: ANDREW NG

 In the 1980s, one better way seemed to be deep learning in neural networks. These systems promised to learn their own rules from scratch, and offered the pleasing symmetry of using brain-inspired mechanics to achieve brain-like function. The strategy called for simulated neurons to be organized into several layers. Give such a system a picture and
  • the first layer of learning will simply notice all the dark and light pixels. 
  • The next layer might realize that some of these pixels form edges; 
  • the next might distinguish between horizontal and vertical lines. 
  • Eventually, a layer might recognize eyes, and might realize that two eyes are usually present in a human face (see 'Facial recognition').

The first deep-learning programs did not perform any better than simpler systems, says Malik. Plus, they were tricky to work with. “Neural nets were always a delicate art to manage. There is some black magic involved,” he says. The networks needed a rich stream of examples to learn from — like a baby gathering information about the world. In the 1980s and 1990s, there was not much digital information available, and it took too long for computers to crunch through what did exist. Applications were rare. One of the few was a technique — developed by LeCun — that is now used by banks to read handwritten cheques.

By the 2000s, however, advocates such as LeCun and his former supervisor, computer scientist Geoffrey Hinton of the University of Toronto in Canada, were convinced that increases in computing power and an explosion of digital data meant that it was time for a renewed push. “We wanted to show the world that these deep neural networks were really useful and could really help,” says George Dahl, a current student of Hinton's.

As a start, Hinton, Dahl and several others tackled the difficult but commercially important task of speech recognition. In 2009, the researchers reported2 that after training on a classic data set — three hours of taped and transcribed speech — their deep-learning neural network had broken the record for accuracy in turning the spoken word into typed text, a record that had not shifted much in a decade with the standard, rules-based approach. The achievement caught the attention of major players in the smartphone market, says Dahl, who took the technique to Microsoft during an internship. “In a couple of years they all switched to deep learning.” For example, the iPhone's voice-activated digital assistant, Siri, relies on deep learning.

Giant leap
When Google adopted deep-learning-based speech recognition in its Android smartphone operating system, it achieved a 25% reduction in word errors. “That's the kind of drop you expect to take ten years to achieve,” says Hinton — a reflection of just how difficult it has been to make progress in this area. “That's like ten breakthroughs all together.

Meanwhile, Ng had convinced Google to let him use its data and computers on what became Google Brain. The project's ability to spot cats was a compelling (but not, on its own, commercially viable) demonstration of unsupervised learning — the most difficult learning task, because the input comes without any explanatory information such as names, titles or categories. But Ng soon became troubled that few researchers outside Google had the tools to work on deep learning. “After many of my talks,” he says, “depressed graduate students would come up to me and say: 'I don't have 1,000 computers lying around, can I even research this?'”

So back at Stanford, Ng started developing bigger, cheaper deep-learning networks using graphics processing units (GPUs) — the super-fast chips developed for home-computer gaming3. Others were doing the same. “For about US$100,000 in hardware, we can build an 11-billion-connection network, with 64 GPUs,” says Ng.

Victorious machine

But winning over computer-vision scientists would take more: they wanted to see gains on standardized tests. Malik remembers that Hinton asked him: “You're a sceptic. What would convince you?” Malik replied that a victory in the internationally renowned ImageNet competition might do the trick.

In that competition, teams train computer programs on a data set of about 1 million images that have each been manually labelled with a category. After training, the programs are tested by getting them to suggest labels for similar images that they have never seen before. They are given five guesses for each test image; if the right answer is not one of those five, the test counts as an error. Past winners had typically erred about 25% of the time. In 2012, Hinton's lab entered the first ever competitor to use deep learning. It had an error rate of just 15% (ref. 4).

Deep learning stomped on everything else,” says LeCun, who was not part of that team. The win landed Hinton a part-time job at Google, and the company used the program to update its Google+ photo-search software in May 2013.

Malik was won over. “In science you have to be swayed by empirical evidence, and this was clear evidence,” he says. Since then, he has adapted the technique to beat the record in another visual-recognition competition5. Many others have followed: in 2013, all entrants to the ImageNet competition used deep learning.


“Over the next few years we'll see a feeding frenzy. Lots of people will jump on the deep-learning bandwagon.”

With triumphs in hand for image and speech recognition, there is now increasing interest in applying deep learning to natural-language understanding — comprehending human discourse well enough to rephrase or answer questions, for example — and to translation from one language to another. Again, these are currently done using hand-coded rules and statistical analysis of known text. The state-of-the-art of such techniques can be seen in software such as Google Translate, which can produce results that are comprehensible (if sometimes comical) but nowhere near as good as a smooth human translation. “Deep learning will have a chance to do something much better than the current practice here,” says crowd-sourcing expert Luis von Ahn, whose company Duolingo, based in Pittsburgh, Pennsylvania, relies on humans, not computers, to translate text. “The one thing everyone agrees on is that it's time to try something different.

Deep science

In the meantime, deep learning has been proving useful for a variety of scientific tasks. “Deep nets are really good at finding patterns in data sets,” says Hinton. In 2012, the pharmaceutical company Merck offered a prize to whoever could beat its best programs for helping to predict useful drug candidates. The task was to trawl through database entries on more than 30,000 small molecules, each of which had thousands of numerical chemical-property descriptors, and to try to predict how each one acted on 15 different target molecules. Dahl and his colleagues won $22,000 with a deep-learning system. “We improved on Merck's baseline by about 15%,” he says.

Biologists and computational researchers including Sebastian Seung of the Massachusetts Institute of Technology in Cambridge are using deep learning to help them to analyse three-dimensional images of brain slices. Such images contain a tangle of lines that represent the connections between neurons; these need to be identified so they can be mapped and counted. In the past, undergraduates have been enlisted to trace out the lines, but automating the process is the only way to deal with the billions of connections that are expected to turn up as such projects continue. Deep learning seems to be the best way to automate. Seung is currently using a deep-learning program to map neurons in a large chunk of the retina, then forwarding the results to be proofread by volunteers in a crowd-sourced online game called EyeWire.

“Deep learning has the property that if you feed it more data, it gets better and better.”

William Stafford Noble, a computer scientist at the University of Washington in Seattle, has used deep learning to teach a program to look at a string of amino acids and predict the structure of the resulting protein — whether various portions will form a helix or a loop, for example, or how easy it will be for a solvent to sneak into gaps in the structure. Noble has so far trained his program on one small data set, and over the coming months he will move on to the Protein Data Bank: a global repository that currently contains nearly 100,000 structures.

For computer scientists, deep learning could earn big profits: Dahl is thinking about start-up opportunities, and LeCun was hired last month to head a new AI department at Facebook. The technique holds the promise of practical success for AI. “Deep learning happens to have the property that if you feed it more data it gets better and better,” notes Ng. “Deep-learning algorithms aren't the only ones like that, but they're arguably the best — certainly the easiest. That's why it has huge promise for the future.

Not all researchers are so committed to the idea. Oren Etzioni, director of the Allen Institute for Artificial Intelligence in Seattle, which launched last September with the aim of developing AI, says he will not be using the brain for inspiration. It's like when we invented flight,” he says; the most successful designs for aeroplanes were not modelled on bird biology. Etzioni's specific goal is to invent a computer that, when given a stack of scanned textbooks, can pass standardized elementary-school science tests (ramping up eventually to pre-university exams). To pass the tests, a computer must be able to read and understand diagrams and text. How the Allen Institute will make that happen is undecided as yet — but for Etzioni, neural networks and deep learning are not at the top of the list.

One competing idea is to rely on a computer that can reason on the basis of inputted facts, rather than trying to learn its own facts from scratch. So it might be programmed with assertions such as 'all girls are people'. Then, when it is presented with a text that mentions a girl, the computer could deduce that the girl in question is a person. Thousands, if not millions, of such facts are required to cover even ordinary, common-sense knowledge about the world. But it is roughly what went into IBM's Watson computer, which famously won a match of the television game show Jeopardy against top human competitors in 2011. Even so, IBM's Watson Solutions has an experimental interest in deep learning for improving pattern recognition, says Rob High, chief technology officer for the company, which is based in Austin, Texas.

Google, too, is hedging its bets. Although its latest advances in picture tagging are based on Hinton's deep-learning networks, it has other departments with a wider remit. In December 2012, it hired futurist Ray Kurzweil to pursue various ways for computers to learn from experience — using techniques including but not limited to deep learning. Last May, Google acquired a quantum computer made by D-Wave in Burnaby, Canada (see Nature 498, 286–288; 2013). This computer holds promise for non-AI tasks such as difficult mathematical computations — although it could, theoretically, be applied to deep learning.

Despite its successes, deep learning is still in its infancy. “It's part of the future,” says Dahl. “In a way it's amazing we've done so much with so little.” And, he adds, “we've barely begun”. Nature505,146–148(09 January 2014)doi:10.1038/505146a
References Le, Q. V. et al. Preprint at http://arxiv.org/abs/1112.6209 (2011). Show context
Mohamed, A. et al. 2011 IEEE Int. Conf. Acoustics Speech Signal Process. http://dx.doi.org/10.1109/ICASSP.2011.5947494 (2011). Show context
Coates, A. et al. J. Machine Learn. Res. Workshop Conf. Proc. 28, 1337–1345 (2013). Show context
Krizhevsky, A., Sutskever, I. & Hinton, G. E. In Advances in Neural Information Processing Systems 25; available at http://go.nature.com/ibace6 Show context
Girshick, R., Donahue, J., Darrell, T. & Malik, J. Preprint at http://arxiv.org/abs/1311.2524 (2013).

martes, 10 de julio de 2012

An expanded genetic alphabet could lead to more easily designed proteins


Floyd E. Romesberg, associate professor at Scripps Research Institute (Credit: The Scripps Research Institute)
Back in 1992, in chapter 15 of Nanosystems, Eric Drexler suggested that it would be easier to design proteins that fold predictably, an important step on the road to advanced nanotechnology (or molecular manufacturing, or atomically precise manufacturing) if additional amino acids beyond the 20 that are genetically coded could be incorporated into proteins. Chemical synthesis of peptides has provided a way to accomplish this for small amounts of short proteins, but to obtain large amounts of long proteins, it would be very convenient to expand the genetic alphabet to encode additional amino acids. This long-standing effort has taken a major step forward with the discovery of how artificial DNA base pairs can be replicated. A hat tip to ScienceDaily for reprinting this Scripps Research Institute news release “Scripps Research Institute study suggests expanding the genetic alphabet may be easier than previously thought“:

A new study led by scientists at The Scripps Research Institute suggests that the replication process for DNA—the genetic instructions for living organisms that is composed of four bases (C, G, A and T)—is more open to unnatural letters than had previously been thought. An expanded “DNA alphabetcould carry more information than natural DNA, potentially coding for a much wider range of molecules and enabling a variety of powerful applications, from precise molecular probes and nanomachines to useful new life forms.

The new study, which appears in the June 3, 2012 issue of Nature Chemical Biology [abstract], solves the mystery of how a previously identified pair of artificial DNA bases can go through the DNA replication process almost as efficiently as the four natural bases.

We now know that the efficient replication of our unnatural base pair isn’t a fluke, and also that the replication process is more flexible than had been assumed,” said Floyd E. Romesberg, associate professor at Scripps Research, principal developer of the new DNA bases, and a senior author of the new study. The Romesberg laboratory collaborated on the new study with the laboratory of co-senior author Andreas Marx at the University of Konstanz in Germany, and the laboratory of Tammy J. Dwyer at the University of San Diego.

Adding to the DNA Alphabet
Romesberg and his lab have been trying to find a way to extend the DNA alphabet since the late 1990s. In 2008, they developed the efficiently replicating bases NaM and 5SICS, which come together as a complementary base pair within the DNA helix, much as, in normal DNA, the base adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G).

The following year, Romesberg and colleagues showed that NaM and 5SICS could be efficiently transcribed into RNA in the lab dish. But these bases’ success in mimicking the functionality of natural bases was a bit mysterious. They had been found simply by screening thousands of synthetic nucleotide-like molecules for the ones that were replicated most efficiently. And it had been clear immediately that their chemical structures lack the ability to form the hydrogen bonds that join natural base pairs in DNA. Such bonds had been thought to be an absolute requirement for successful DNA replication-—a process in which a large enzyme, DNA polymerase, moves along a single, unwrapped DNA strand and stitches together the opposing strand, one complementary base at a time.

An early structural study of a very similar base pair in double-helix DNA added to Romesberg’s concerns. The data strongly suggested that NaM and 5SICS do not even approximate the edge-to-edge geometry of natural base pairs—termed the Watson-Crick geometry, after the co-discoverers of the DNA double-helix. Instead, they join in a looser, overlapping, “intercalated” fashion. “Their pairing resembles a ‘mispair,’ such as two identical bases together, which normally wouldn’t be recognized as a valid base pair by the DNA polymerase,” said Denis Malyshev, a graduate student in Romesberg’s lab who was lead author along with Karin Betz of Marx’s lab.

Yet in test after test, the NaM-5SICS pair was efficiently replicable. “We wondered whether we were somehow tricking the DNA polymerase into recognizing it,” said Romesberg. “I didn’t want to pursue the development of applications until we had a clearer picture of what was going on during replication.

Edge to Edge
To get that clearer picture, Romesberg and his lab turned to Dwyer’s and Marx’s laboratories, which have expertise in finding the atomic structures of DNA in complex with DNA polymerase. Their structural data showed plainly that the NaM-5SICS pair maintain an abnormal, intercalated structure within double-helix DNA—but remarkably adopt the normal, edge-to-edge, “Watson-Crick” positioning when gripped by the polymerase during the crucial moments of DNA replication.

The DNA polymerase apparently induces this unnatural base pair to form a structure that’s virtually indistinguishable from that of a natural base pair,” said Malyshev.

NaM and 5SICS, lacking hydrogen bonds, are held together in the DNA double-helix by “hydrophobic” forces, which cause certain molecular structures (like those found in oil) to be repelled by water molecules, and thus to cling together in a watery medium. “It’s very possible that these hydrophobic forces have characteristics that enable the flexibility and thus the replicability of the NaM-5SICS base pair,” said Romesberg. “Certainly if their aberrant structure in the double helix were held together by more rigid covalent bonds, they wouldn’t have been able to pop into the correct structure during DNA replication.

An Arbitrary Choice?
The finding suggests that NaM-5SICS and potentially other, hydrophobically bound base pairs could some day be used to extend the DNA alphabet. It also hints that Evolution’s choice of the existing four-letter DNA alphabet—on this planet—may have been somewhat arbitrary.It seems that life could have been based on many other genetic systems,” said Romesberg.

He and his laboratory colleagues are now trying to optimize the basic functionality of NaM and 5SICS, and to show that these new bases can work alongside natural bases in the DNA of a living cell.

If we can get this new base pair to replicate with high efficiency and fidelity in vivo, we’ll have a semi-synthetic organism,” Romesberg said. “The things that one could do with that are pretty mind blowing.

Now that these scientists have demonstrated that DNA replication is far more flexible than had been thought, it will be fascinating to see what researchers do to expand the genetic alphabet, what additional amino acids they choose to incorporate into what proteins, and what this expansion of the set of amino acids composing proteins means for predicable protein folding and for artificial protein molecular machines.
—James Lewis, PhD

domingo, 10 de junio de 2012

Researchers Watch Tiny Living Machines Self-Assemble

ORIGINAL: Science Daily

Vallée-Bélisle and Michnick have developed a new approach to visualize how proteins assemble, which may also significantly aid our understanding of diseases such as Alzheimer’s and Parkinson’s, which are caused by errors in assembly. Here shown are two different assembly stages (purple and red) of the protein ubiquitin and the fluorescent probe used to visualize these stage (tryptophan: see yellow). (Credit: Peter Allen)
ScienceDaily (June 10, 2012) — Enabling bioengineers to design new molecular machines for nanotechnology applications is one of the possible outcomes of a study by University of Montreal researchers that was published in Nature Structural and Molecular Biology June 10. The scientists have developed a new approach to visualize how proteins assemble, which may also significantly aid our understanding of diseases such as Alzheimer's and Parkinson's, which are caused by errors in assembly. 

"In order to survive, all creatures, from bacteria to humans, monitor and transform their environments using small protein nanomachines made of thousands of atoms," explained the senior author of the study, Prof. Stephen Michnick of the university's department of biochemistry. "For example, in our sinuses, there are complex receptor proteins that are activated in the presence of different odor molecules. Some of those scents warn us of danger; others tell us that food is nearby." Proteins are made of long linear chains of amino acids, which have evolved over millions of years to self-assemble extremely rapidly - often within thousandths of a split second -- into a working nanomachine. "One of the main challenges for biochemists is to understand how these linear chains assemble into their correct structure given an astronomically large number of other possible forms," Michnick said.

"To understand how a protein goes from a linear chain to a unique assembled structure, we need to capture snapshots of its shape at each stage of assembly" said Dr. Alexis Vallée-Bélisle, first author of the study. "The problem is that each step exists for a fleetingly short time and no available technique enables us to obtain precise structural information on these states within such a small time frame. We developed a strategy to monitor protein assembly by integrating fluorescent probes throughout the linear protein chain so that we could detect the structure of each stage of protein assembly, step by step to its final structure."

The protein assembly process is not the end of its journey, as a protein can change, through chemical modifications or with age, to take on different forms and functions. "Understanding how a protein goes from being one thing to becoming another is the first step towards understanding and designing protein nanomachines for biotechnologies such as medical and environmental diagnostic sensors, drug synthesis of delivery," Vallée-Bélisle said.

This research was supported by the Natural Sciences and Engineering Research Council of Canada and Le fond de recherché du Québec, Nature et Technologie. The article, "Visualizing transient protein folding intermediates by tryptophan scanning mutagenesis," published in Nature Structural & Molecular Biology, was coauthored by Alexis Vallée-Bélisle and Stephen W. Michnick of the Département de Biochimie de l'Université de Montréal. The University of Montreal is known officially as Université de Montréal.

lunes, 7 de mayo de 2012

Synthetic Biological Life

ORIGINAL: HPlusMagazine
By: Laura E. Bratton, MD, Rodney Shackelford, DO, Ph.D.
May 3, 2012

The idea of producing artificial or synthetic life has long fascinated mankind and from ancient times many human and animal-imitating “automata” or self-operating machines have been created for entertainment, instructional, and sometimes religious purposes. The creation of actual synthetic biological life only became possible with the discovery of the structure of DNA, the genetic code, and the development of the basic tools of molecular biology, such as the ability to isolate, sequence, and join different DNA sequences. Especially important has been the recently developed ability to artificially synthesize relatively long DNA molecules with designed sequences. Although the creation of completely synthetic biological life was first accomplished in 2010, the field is already yielding significant information concerning the core gene groups or genetic “chassis” indispensible for life and how these gene products (proteins, RNAs, and lipids) function as an integrated unit. With the identification of these chassis, exogenous natural or synthetic gene sequences can be integrated into organisms designed for specific purposes and applications.

The first genetically engineered organism was created in 1973 when a naturally occurring DNA sequence was transferred into and expressed in a bacterium, conferring antibiotic resistance. The first organism to actually have a synthetic (or man-made “added”) biochemical pathway was created in 2003, when an E. coli was artificially created with a new genetic code and amino acid synthesizing enzymes. The engineered bacterium could synthesize and incorporate an amino acid (O-methyl-L-tyrosine) that does not normally occur in nature into proteins, increasing the number of amino acids used in virtually all life forms from twenty to twenty-one amino acids. Thus a new, human-designed functioning genetic chassis and genetic code was placed into a microorganism.

In 2010, after some fifteen years of intense research effort, the first entirely synthetic organism was created with a genome entirely synthesized “out of four bottles” i.e., chemically synthesized from the four DNA bases; thymine, cytosine, guanine, and adenine. The organism was partially based on M. mycoides, a genetically simple microorganism containing roughly 480 protein-encoding genes and a genome size of 1.08 million DNA base pairs – in comparison the human genome has roughly 20,500 genes over three billion DNA base pairs. The synthetic genome was chemically synthesized in 80-90 base units and slowly assembled into “DNA cassettes”, verified by sequencing, and assembled into a circular genome. To insure that no natural DNA contaminated the synthetic DNA “watermark” sequences were inserted into synthetic genome to differentiate it from the natural M. mycoides genome. Additionally, antibiotic resistance genes were added and a disease-inducing gene was removed from the synthetic genome. The resulting genome was place in an empty M. capricolum cell (i.e., without a nucleus) and the resulting synthetic life from was able to grow in culture indefinitely. Since 2010 this synthetic organism has been useful in identifying the “minimal genome” required for life – about 380 of the 480 protein-encoding genes. Additionally, comparison of the synthetic organism to similar naturally occurring organisms (Mycoplasmas), allowed the identification of gene groups involved in cellular processes such as 
  • information storage, 
  • metabolism, 
  • energy production and conversion, and 
  • cell membrane biogenesis. 
Identification of these gene sets is an important first step designing synthetic life that can perform specific functions.

Although a significant first step in the creation of synthetic biological life, this initial work met with extensive criticism. The researchers who made synthetic life were accused of “playing God” and possibly opening up a new technology that would allow the creation of “biological super weapons”. The later objection has some validity, as existing DNA synthesis and end-joining technology could allow the synthesis of fully infective polio or smallpox viruses. Other researchers pointed out that the new synthetic organism was a nearly one-to-one copy of a naturally occurring organism and for it to grow the synthetic genome had to be placed into a naturally occurring Mycoplasma that had its nucleus removed. Thus, other than the DNA being artificially synthesized, there was relatively little that was actually new about the organism. The creators of the new organism pointed out that this is a first step of many and “creating life from scratch” will come later.

Currently the immediate focus in synthetic biological life research is to use simple synthetic organisms to define the “minimal genome”, or the smallest set of genes required to support life and identify the components and functions of “biological gene-chassis” and find ways to modify these chassis. Specific applications include the creation of synthetic organisms that can: 
  1. efficiently produce pharmaceuticals and vaccines that are otherwise difficult and expensive to produce, 
  2. efficiently produce hydrocarbon biofuels (replacing oil, coal, etc.), and 
  3. be useful as plant feedstock in agriculture, lowering the need for increasingly expensive petroleum-based fertilizers. 
An example of such an application has been inserting the enzymes for artemisinic acid synthesis into baker’s yeast. Artemisinic acid is the chemical precursor anti-malarial drug artmisinin, a drug that is currently extracted from the sweet wormwood plant at high cost, reducing the drugs availability in poorer countries. Once the enzymatic pathway is in place and efficiently working, the drug could be produced cheaply in large amounts through a process resembling brewing beer. Several of these projects are being researched at Synthetic Genomics, a new biotechnology company specializing in the creation of synthetic life for specific applications.

Not surprisingly the creation of synthetic animal life is more complex and difficult than for simpler microorganisms. However, a round worm (C. elegans) was created that carried an extensively expanded genetic code and protein synthesis pathways, allowing the incorporation of multiple novel (or “unnatural”) amino acids into the animal’s proteins. These protein modifications would facilitate the study of protein localization and interactions within a living animal. Additionally, modified proteins could be designed for specific purposes, such as protein-based drugs with very long half-lives due to novel amino acids that inhibit normal cellular protein degradation.

Although difficult, our present molecular biology technology could allow the creation of more complex organisms, including fungi and even animals. The present challenges in creating synthetic life include the following:
  1. Create synthetic life “from scratch” without the need to largely copy existing life forms.
  2. Improve on our ability to design and integrate molecular pathways within synthetic life.
  3. Create a strategy or “algorithm” to for the efficient creation of synthetic life forms.
  4. Create policies and rules to prevent the creation of synthetic life forms that may be harmful, such as human pathogens (smallpox, virulent influenza viral types, etc.).
With time these goals could be achieved and the technology to accomplish these goals is largely in place.

In the more distant future synthetic biology could allow the extensive modification of existing genomes and even the creation of entirely new genomes and species. While this is the goal of many Transhumanists, one hopes that if and when such technology exists, the human race has the intelligence to apply such technology with wisdom.

Dr. Shackelford is an Assistant Professor of Clinical Pathology at Tulane Medical Center. He has a DO degree from Des Moines University of Osteopathic Medicine and a Ph.D. in molecular pathology from Duke University. His areas of research include DNA repair, molecular mechanisms of carcinogenesis, and cell division.

martes, 31 de enero de 2012

Folding research recruits unconventional help

Michael Gross is a science writer based at Oxford. He can be contacted via his web page at www.michaelgross.co.uk

Summary
A denatured protein chain can find its well-ordered three-dimensional structure, the native state, in under a second, using only the information contained in the sequence. For researchers, however, the prediction of structures from sequences is a hard problem, so they are now recruiting all the help they can get, including idle computers and game consoles, game players, and little hints from evolution. Michael Gross reports.

MAIN TEXT
Protein folding is one of the miracles of nature that human technology finds quite difficult to follow. Ever since the classic ribonuclease A experiments of Christian Anfinsen in the 1960s it has been clear that the amino acid sequence of a polypeptide chain determines the unique three-dimensional folded conformation it will adopt under physiological conditions. Even though we now know that some proteins remain intrinsically disordered, and some may ‘fold around’ a ligand, it is still true that a protein's structure, and hence its function, is somehow encoded in the sequence of the amino acids. As the theoretician Cyrus Levinthal pointed out early on, there is an astronomical number of wrong conformations which the chain cannot possibly try out in a reasonable time, so there must be mechanisms that allow polypeptide chains to find the native state encoded in their sequence. Folding researchers have elucidated some of these mechanisms in the last decades, but so far haven't been able to decipher the code and therefore aren't generally able to read a sequence and predict what shape it will adopt.

An early approach to the problem was to build bigger computers dedicated to simulations of folding, a development that culminated in the development of IBM's Blue Gene, first announced in 1999 with the explicit target of tackling protein folding, but it didn't crack the prediction problem once and for all. By the turn of the millennium, computer simulations of the protein movements could only just cover a microsecond, while it was known from experimental studies that the relevant folding reactions happened on the millisecond timescale.

Folding at home


Based on that shortfall of computing power and the increasing availability of PCs connected via the internet, the group of Vijay Pande at Stanford University developed a distributed computing program called Folding@home, using the now widely adopted practice of chopping a problem into small parcels and farming them out to many computers that are online but idle (as, for instance, most computers in universities are, most of the time).

Since 2006, the lab also offers a version of the program that runs on games consoles, which enabled a dramatic increase in the processing power accessible to the program. In September 2007, the program achieved a consistent level above one petaFLOPS (1015 floating-point operations per second), as the first computing system of any kind to do so (the fastest supercomputer at the time was Blue Gene with just over a quarter of that power). In November 2011, Folding@home passed the milestone of six petaFLOPS.
Folding puzzle: The rearrangement of a linear polymer into a compact three-dimensional shape proceeds autonomously in nature but is still puzzling biochemists. (Photo: Michael Gross.)
The simulation of protein dynamics by the Folding@home software is based on the observation that protein chains populate certain free energy minima for a period of time, and then quickly move on to another minimum. Pande's group uses so-called Markov state models to find connections between these minima and map the likelihood of transitions. In 2010, the group used this approach combined with the distributed computing power of Folding@home to simulate the folding of the 39-residue protein NTL9, which takes about 1.5 milliseconds (J. Am. Chem. Soc. (2010), 132, 1526–1528). The time span of this simulation was a thousand times longer than that achieved by other methods.

Pande's group uses this approach to address a range of biomedically relevant issues around protein folding, including diseases that involve misfolded proteins (such as Alzheimer's, Parkinson's and Huntington's disease), the fundamental questions of protein folding mechanisms, and the prediction of unknown structures from sequences, which feed into commercial drug design. “Through the Folding@home distributed computing project, we have been able to muster an unparalleled computational resource for studying protein folding, allowing us to study complex systems on timescales thousands of times longer than would otherwise be possible,” says Pande. The group makes all datasets generated from the research available on request.

Folding game


Meanwhile, the group of David Baker at the University of Washington at Seattle had developed an algorithm called Rosetta for the prediction of small protein structures, and also set up distributed computing (Rosetta@home) to provide additional computing power for the prediction work. Participants who have this program installed on their computers can watch how the protein chain gradually finds its native conformation. It so happened that some participants watched the process and spotted possible arrangements that would improve the energy-efficient packing, but weren't able to interact with the program, finding themselves in the situation of someone watching a game show and shouting at their TV set.
Folding fun: Tens of thousands of online game players around the world have contributed to folding research via the game Foldit, developed at the University of Washington at Seattle. (Photo: Mohini Patel Glanz.)