Mostrando entradas con la etiqueta XML. Mostrar todas las entradas
Mostrando entradas con la etiqueta XML. Mostrar todas las entradas

lunes, 16 de noviembre de 2015

PLOS and DBpedia – an experiment towards Linked Data



Editor’s Note: This article is coauthored by Bob Kasenchak, Director of Business Development/Taxonomist at Access Innovations.

PLOS publishes articles covering a huge range of disciplines. This was a key factor in PLOS deciding to develop its own thesaurus – currently with 10,767 Subject Area terms for classifying the content.

We wondered whether matching software could establish relationships between PLOS Subject Areas and corresponding terms in external datasets. These relationships could enable links between data resources and expose PLOS content to a wider audience. So we set out to see if we could populate a field for each term in the PLOS thesaurus with a link to an external resource that describes—or, is “the same as”—the concept in the thesaurus. If so, we could:

• Provide links between PLOS Subject Area pages and external resources
• Import definitions to the PLOS thesaurus from matching external resources

For example, adding Linked Data URIs to the Subject Areas would facilitate making the PLOS thesaurus available as part of the Semantic Web of linked vocabularies.

We decided to use DBpedia for this trial for two reasons:

Firstly, as stated by DBpediaThe DBpedia knowledge base is served as Linked Data on the Web. As DBpedia defines Linked Data URIs for millions of concepts, various data providers have started to set RDF links from their data sets to DBpedia, making DBpedia one of the central interlinking-hubs of the emerging Web of Data.

Figure 1: Linked Open Data Cloud with PLOS shown linking to DBpedia – the concept behind this project.
Secondly, DBpedia is constantly (albeit slowly) updated based on frequently-used Wikipedia pages; so has a method to stay current, and a way to add content to DBpedia pages, providing inbound links—so people can link (directly or indirectly) to PLOS Subject Area Landing Pages via DBpedia.

Figure 2: ‘Cognitive psychology’ pages in PLOS and DBpedia 
Which matching software to trial?
We considered two possibilities: Silk and Spotlight

  • The Silk method might have allowed more granular, specific, and accurate queries, but it would have required us to learn a new query language. 
  • Spotlight, on the other hand, is executable by a programmer via API and required little effort to run, re-run, and check results; it took only a matter of minutes to get results from a list of terms to match. 

So we decided to use Spotlight for this trial.

Which sector of the thesaurus to target?
We chose the Psychology division of 119 terms (see Appendix) as a good starting point because it provides a reasonable number of test terms so that trends could emerge, and a range of technical terms (e.g. Neuropsychology) as well as general-language terms (e.g. Attention) to test the matching software.

Methods:

Figure 3: Work flow.
Step 1: We created the External Link and Synopsis DBpedia fields in the MAIstro Thesaurus Master application to store the identified external URIs and definitions. The External Link field accommodates the corresponding external URI, and the Synopsis DBpedia field houses the definition – “dbo:abstract” in DBpedia.
Step 2: Matching DBpedia concepts with PLOS Subject Areas using Spotlight:

  • Phase 1: For the starting set of terms we chose Psychology (a Tier 2 term) and the 21 Narrower Terms that sit in Tier 3 immediately beneath Psychology (listed in Appendix ).
  • Phase 2: For Phase 2 we included the remaining 98 terms from Tier 4 and deeper beneath Psychology (listed in Appendix ).
Step 3: Importing External Link/Synopsis DBpedia to PLOS Thesaurus: Once a list of approved matching PLOS-term-to-DBpedia-page correspondences was established, another quick DBpedia Spotlight query provided the corresponding Definitions. Access Innovations populated the fields by loading the links and definitions to the corresponding term records. For the “Cognitive psychology” example these are:

Synopsis DBpedia: Cognitive psychology is the study of mental processes such as “attention, language use, memory, perception, problem solving, creativity, and thinking.” Much of the work derived from cognitive psychology has been integrated into various other modern disciplines of psychological study including

  • educational psychology, 
  • social psychology, 
  • personality psychology, 
  • abnormal psychology, 
  • developmental psychology, and 
  • economics.
How did it go?
The table shows the distribution of results for the 119 Subject Areas in the Psychology branch of the PLOS thesaurus:
Add caption
Thus a total of 96 matches could be found by any method (80.7% of terms – top three rows of the Table). Of these, 86 terms (72.3% of terms) were matched as one of the top 5 Spotlight hits (top two rows of the Table), as compared to 71 matches (59.7% of terms) being identified correctly and directly by Spotlight as the top hit (top row of the Table).

Figure 4 shows the two added fields “Synopsis DBpedia” and “External Link” in MAIstro, for “Cognitive Psychology”.

Figure 4: Addition of Synopsis DBpedia and External Link fields to MAIstro.
Conclusions:
We had set out to establish whether matching software could define relationships between PLOS thesaurus terms and corresponding terms in external datasets. We used the Psychology division of the PLOS thesaurus as our test vocabulary, Spotlight as our matching software, and DBpedia as our target external dataset.

We found that unambiguous suitable matches were identified for 59.7% of terms. Expressed another way, mismatches were identified as the top hit for 35 cases (29.4% of terms) which is a high burden of inaccuracy. This is too low a quality outcome for us to consider adopting Spotlight suggestions without editorial review.

As well as those terms that were matched as a top hit, a further 12.6% of terms (or 31% of the terms not successfully matched as a top hit) had a good match in Spotlight hit positions 2-5. So Spotlight successfully matched 72.3% of terms within in the top 5 Spotlight matches.

Having the Spotlight hit list for each term did bring efficiency to finding the correspondences. Both the “hits” and the “misses” were straightforward to identify. As an aid to the manual establishment of these links Spotlight is extremely useful.

Stability of DBpedia: We noticed that the dbo:abstract content in DBpedia is not stable. It would be an enhancement to append the Synopsis DBpedia field contents with URI and date stamp as a rudimentary versioning/quality control measure.

Can we improve on Spotlight? Possibly. We wouldn’t be comfortable with any scheme that linked PLOS concepts to the world of Linked Data sources without editorial quality control. But we suspect that a more sophisticated matching tool might promote those hits that fell within Spotlight matches 2-5 to the top hit, and would find some of the 8.4% of terms which were found manually but which Spotlight did not suggest in the top 5 hits at all. We hope to invest some effort in evaluating Silk, and establishing whether or not any other contenders are emerging.

Introducing PLOS Subject Area URIs into DBpedia page: This was explored and it seemed likely that the route to achieve this would be to add the PLOS URI first to the corresponding Wikipedia page, in the “External Links” section.

Figure 5: The External Links section of Wikipedia: Cognitive psychology
As DBpedia (slowly) crawls through all associated Wikipedia pages, eventually the new PLOS link would be added to the DBpedia entry for each page.

To demonstrate this methodology, we added a backlink to the corresponding PLOS Subject Area page in the Wikipedia article shown above (Cognitive psychology) as well as all 21 Tier 3 Psychology terms.

Figure 6: External Links at Wikipedia: Cognitive psychology showing link back to the corresponding PLOS Subject Area page
Were DBpedia to re-crawl this page, the link to the PLOS page would be added to DBpedia’s corresponding page as well.

However, Wikipedia questioned the value of the PLOS backlinks (“link spam”) and their appropriateness to the “External Links” field in the various Wikipedia pages. A Wiki administrator can deem them inappropriate and remove them from Wikipedia (as has happened for some if not all of them by the time you read this).

We believe the solution is to publish the PLOS thesaurus as Linked Open Data (in either SKOS or OWL format(s)) and assert the link to the published vocabulary from DBpedia (using the field owl:sameAs instead of dbo:wikiPageExternalLink). We are looking into the feasibility and mechanics of this.

Once the PLOS thesaurus is published in this way, the most likely candidate for interlinking data would be to use the SILK Linked Data Integration Framework and we look forward to exploring that possibility.

Appendix: The Psychology division of the PLOS thesaurus. LD_POC_blog.appendix


ORIGINAL: PlOS
 November 10, 2015

martes, 25 de noviembre de 2014

Open qPCR: DNA Diagnostics for Everyone

Turning DNA into data just became affordable, and biohacking will never be the same. Meet our open source Real-Time PCR Thermocycler.



Real-Time PCR is a powerful technology. It can detect foodborne contaminants like E. Coli and Listeria, as well identify fraudulently labeled food products. It's used to diagnose infections such as HIV and Malaria, and we've created a reward level to donate an Open qPCR to Ebola clinics in Western Africa.

Real-Time PCR can also identify genetic mutations that increase our likelihood of cancers, such as those within BRCA1. It's a fundamental tool of biological research, and has literally thousands of other applications too long to list here.

However with machines costing $20,000 and up, Real-Time PCR (also known as qPCR) is simply unaffordable where it is needed most. We aim to change that with Open qPCR, a system available at a fraction of their typical cost.

Four years ago we Kickstarted and successfully delivered OpenPCR, the world's first open source thermocycler. Now we're taking it to the next level with a machine that not only copies the DNA but converts it into data.

Open qPCR is disrupting DNA diagnostics.
Our reason for starting this project is simple: we want to make this essential technology available to everyone, including doctors in developing countries, students in high school and university labs, companies in the food supply chain, and biohackers who are developing some of the most innovative synthetic biology applications.


Over the past two years, we've designed and prototyped an open-source instrument that can perform the above mentioned tests, but costs less than a tenth the cost of other commercially available systems. Now, we're ready to share it with the world and see what others can do with it.

DNA Diagnostics on your Desktop
To create a truly cost-effective diagnostic system, we realized users would also need affordable access to the necessary reagents. That's why we created our own low-cost Real-Time PCR MasterMix incorporating our open source fluorescent dye.

We also focused on making the software as user-friendly as possible. While there's powerful functionality available for scientists, the machine also interprets the data and presents clear positive/negative results to end users.


Open qPCR web UI showing detection interpretation. Results are labeled indeterminate if replicates vary or control reactions fail.
What Open qPCR Can Do
The Hardware
Open qPCR's heatblock holds 16 reaction wells capable of holding either 100 or 200 uL tubes/strips. The heat block ramps between 0 and 100 degrees C at 5 C/s, allowing many PCR protocols to be completed in as little as 30 minutes.


Open qPCR's well block supports 16 samples

Fluorescence detection is performed with a solid-state LED excitation (470-490 nm) and photodiode detection system, utilizing a set of optical filters to create sharp cutoffs.

Open qPCR is available in both single-channel and dual-channel versions. Both versions can detect many green fluorophores like FAM in the 510-545 nm range. The dual-channel version can additionally detect HEX/VIC/JOE fluorophores in the 560-610 nm region.


The optical system detects the DNA in the right tube, but the left tube without DNA fails to fluoresce. The single-channel detection system is shown in this image.

An 800x480 pixel capacitive touchscreen allows easy control of the machine, while USB, ethernet, and wifi interfaces let the user control the machine via a web browser or other networked systems. The machine is controlled by an embedded BeagleBone Black microcomputer, running embedded Linux at 1 GHz and sporting 4 GB of flash.


Open qPCR uses ethernet, wifi, and USB to communicate with the outside world

Open qPCR is compatible with power systems around the world (110 - 240 V, 47-63 Hz), and complies with all CE and FCC requirements. In normal cycling operation the machine requires 250 W, but in a special low-power mode, isothermal DNA detection can be performed with as little as 45 W, allowing field use with solar/battery power.

The hardware design, including BOM, SolidWorks files, and Eagle CAD will all be released as open source when the machine ships, and are available sooner as an early access reward.

Tech Specs

The Software
We've developed an open and intuitive user interface using modern web technologies such as HTML5 and JavaScript. Open qPCR runs an embedded web server, which lets the machine be operated in a platform-independent way from nearby computers connected via ethernet or wifi. A USB connectivity mode is also supported for non-networked environments.

Open qPCR provides a powerful protocol editor, which supports basic cycle and step editing, as well as advanced features such as controlled ramp rates and auto-delta/touchdown steps. Data collection may be triggered on each step and/or ramp, and a final hold temperature is supported.

A visual plate layout editor allows the assignment of samples, targets, standards, and controls to each well, as well as viewing fluorescence curves on a per-well basis.

The software supports amplification curve Ct thresholds, presence/absence detection, melt curve analysis, and relative quantification. Software support for absolute quantification may be ready by the shipping date, or will otherwise be made available shortly thereafter as a free downloadable update. It's a bit tricky to do absolute quantification with a 16 well block, but we've got some ideas. Data may also be exported in RDML format, allowing for more advanced analysis in other software.

All software including real-time control, web user interface, and scientific analysis code will be released as open source when the machine ships.

Real-Time PCR Reagents and Starter Packs
Real-Time PCR reagents have a reputation for being expensive, which is why we're creating our own to coincide with the low cost Open qPCR. We've created our own low-cost PCR MasterMix, and are currently synthesizing Chai Green, an IP-free fluorescent intercalating dye. The composition of these reagents and structure of Chai Green will be released as open source.

We'll be including starter packs of these reagents with all Open qPCR machine rewards so you can evaluate these low-cost reagents yourselves. The starter packs will include enough reagents for 50 PCR reactions.


Of course you may also use existing dyes or probes with fluorophores like FAM with your Open qPCR.

Our Story
Four years ago, in 2010, Josh Perfetto and Tito Jankowski started a Kickstarter campaign for OpenPCR, the world's first open source PCR thermocycler. Josh and Tito successfully fulfilled the Kickstarter, and today approximately 800 OpenPCRs are used around the globe.

Though revolutionary at the time, OpenPCR, like all endpoint PCR thermocyclers, was a relatively basic machine. It simply heats and cools a piece of metal in a precise manner. This thermal cycling facilitates the PCR reaction which selectively amplifies DNA, enabling the many applications of PCR described on this page. However turning this amplified DNA into useful information requires downstream laboratory processes which are too laborious, costly, and error-prone for most diagnostic uses.

This provided the impetus to create Open qPCR, a machine that directly turns DNA into actionable information. In 2012 Josh assembled a team of electrical, mechanical, optical, and software engineers to couple the thermal cycling accuracy of OpenPCR with an optical detection system, powerful modern processor, and modern ethernet, wifi, web, and touch screen interfaces. And unlike OpenPCR, which was a build-it-yourself DIY kit, Open qPCR is a professionally assembled ready-to-use system, though it stays true to its DIYbio roots.

It was a long haul, but we're excited to bring it to you.

The Biohacker Kit
To perform tests with Open qPCR, beyond the machine and provided reagents, you'll need a set of PCR Primers specific for your application and control DNA to have confidence in your results. We will be creating an open source database of PCR assays (the primers, probes, protocols, and controls needed to perform PCR) in 2015, and early beta access to that is available as a KickStarter reward.

However as a first step towards that vision, we're offering a Biohacker Kit available separately or as an add-on to an Open qPCR reward. The Biohacker Kit ships internationally and contains:

  • PCR primer and control library for 5 tests: one pertaining to fraudulently labeled food, another to a gene associated with athletic ability (ACTN3), and three other tests which will be decided on by backers
  • Reagents for 50 DNA extractions
  • Real-Time PCR reagents for 100 reactions
  • A laboratory pipette
  • Two boxes of pipette tips
  • PCR strips
  • Sterile cotton swabs for collecting DNA
With our machine and this kit, you'll have everything you need to perform these tests right out of the box.

What's in the Biohacker Kit
Production Schedule
Over the past two years we've built and tested countless prototypes. We now have a working system, and have suppliers and manufacturers lined up to build our machine. We're ready to get Open qPCR into production, but need your help to finance the first manufacturing run.

We plan to have the first batch of units to fulfill the Kickstarter rewards in March 2015, and the second batch ready in April. The expected delivery date for your batch is shown in the reward section when you pledge.

We are continuing to refine the software, and are on-track for basic informatics (amplification curves with Ct values, melt curves, and presence/absence detection) to be fully implemented by our shipping date as well as relative quantification. Support for absolute quantification is on our development roadmap, but may or may not be ready by the shipping date (it's somewhat tricky with our 16 well block). If it is not ready in time, it will be available as a free downloadable software update shortly thereafter.

Progression of Open qPCR's circuit board

Open qPCR has not yet been submitted to nor cleared by the US FDA for medical uses. Thus within the US and EU it is sold for research, educational, and recreational uses only. We will ship Open qPCR globally but you are responsible for ensuring your intended uses are compatible with the regulations of your country.

Add Ons
The T-Shirt, Coffee Mug, and BioHacker Kit rewards may all be added to the higher reward levels by adding the cost of these rewards to your total pledge. If you are shipping outside the US, please add a $15 shipping fee for T-Shirt or Coffee Mug add-ons, and a $60 shipping fee for BioHacker Kit add-ons. You will be able to specify the add-ons you are requesting in a post-campaign survey.

A reagent-only version of the Biohacker kit is also available for $150, plus $40 for shipping outside the US. This kit includes the DNA extraction buffer, Real-Time PCR MasterMix, primer library, and control DNA of the Biohacker kit, but does not include the pipette, tips, cotton swabs, PCR tubes, or eppendorf tubes.

About Us
Chai Biotechnologies is a biotech startup focussed on making biotechnology more broadly accessible. Chai is based in Santa Clara, California, and was founded by Josh Perfetto, a former co-creator of the OpenPCR open-source thermocycler project KickStarted in 2010. Chai is a team of 8 including experts in molecular biology, mechanical and electrical engineering, software development, and user interface design.


We have worked to remove as much risk as possible before launching this campaign. We have created a working prototype, sourced all key components from suppliers, and identified a local manufacturer for our initial batch, which will help us resolve manufacturing issues quickly. Additionally we have successfully delivered on an earlier Kickstarter project OpenPCR, which was a simpler predecessor to this machine.

While we're happy with our preparations to date, there is always the risk of unforeseen challenges when manufacturing physical products. We will keep you advised of our progress and any unexpected challenges, and will work our hardest to meet the delivery schedule.

ORIGINAL: Kickstarter

domingo, 19 de enero de 2014

How to find an appropriate research data repository?

As more and more funders and journals adopt data policies that require researchers to deposit underlying research data in a data repository, the question how to choose a repository becomes more and more important. Heinz Pampel is one of the people behind re3data.org, an Open Science tool that helps researchers to easily identify a suitable repository for their data and thus comply to requirements set out in data policies.

Background

The debate on open access to research data is gaining relevance. This February, the federal agencies in the U.S. have been told by the Office of Science and Technology Policy (OSTP) to maximize access to data from publicly funded research. In June, the G8 science ministers published a set of principles for open scientific research data. The ministers declared that, if possible, „publicly funded scientific research data should be open“. And already last year, the European Commission announced a pilot framework in Horizon 2020, the coming EU framework programme for research and innovation, to promote open access to research data.

Although scientists agree with the potential benefit of data sharing for the scientific progress, the majority is reserved when it comes to practical implementations. One reason for the reluctance is a lack of reliable “systems that make it quick and easy to share data” (Tenopir et al. 2011).

The current landscape of data repositories is heterogeneous. Some initiatives like the Data Seal of Approval (DSA) and the World Data System (WDS) are working on the standardization of data repositories. And there are already certification and auditing procedures for data repositories. Two examples are the DIN 31644 and the ISO 16363 standards. But these standards are not widely used yet. Research data repositories and their services are mostly characterized by the scientific discipline in which they work. They store a wide variety of file formats under different conditions for access and reuse. In many cases it is difficult for researchers to find an appropriate repository for the storage of their data. To overcome these shortcomings we started re3data.orgRegistry of Research Data Repositories.

re3data.org – Registry of Research Data Repositories

Launched in 2012, re3data.org provides an overview of existing research data repositories. In September 2013, re3data.org lists 600 research data repositories, 400 of these are described in detail by a comprehensive vocabulary. The registry covers data repositories from all academic disciplines.

In re3data.org researchers can easily see the terms of access and use of each data repositories and other characteristics. Information icons help researchers to easily identify an adequate repository for the storage and reuse of their data.


Aspects of a Research Data Repository with the corresponding icons used in re3data.org.

re3data.org covers the following aspects of a research data repository:
  • general information (e.g. short description of the repository, content types, keywords),
  • responsibilities (e.g. institutions responsible for funding, content or technical issues),
  • policies (e.g. guidelines and policies of the repository),
  • legal aspects (e.g. licenses of the database and datasets),
  • technical standards (e.g. APIs, versioning of datasets, software of the repository),
  • quality standards (e.g. certificates, audit processes).
The re3data.org portal offers two search possibilities:
  • (1) free text search through a simple search box, and 
  • (2) filters for more specific searches. 
In the list of results each record includes the name of the repository, the subjects covered, a brief description of the content and a set of icons visualizing key properties of the repository. A comprehensive view of the descriptive record of the repository can be obtained by clicking on the name of the repository in the search results. It is also possible to simply browse through the list of indexed data repositories.

Example screenshot of search results for geosciences data repositories using persistent identifiers.

Operators of data repositories can suggest their infrastructures to be listed in re3data.org by filling in an online application form. A repository is indexed when the minimum requirements for inclusion in re3data.org are met. These requirements are described in the re3data.org vocabulary. The project team reviews each repository and reviewed repositories are identified by a green check mark.

The project cooperates with other Open Science initiatives like
Some publishers already refer to re3data.org in their Editorial Policies as a tool for the identification of suitable data repositories.
Next Steps
In the upcoming project phase the focus will be on improving usability and implementing new features. Among other things, the dialog with repositories operators will be supported by a workflow system. Beyond the development of the registry, the project will promote the standardization of research data repositories.

re3data.org is funded by the German Research Foundation (DFG). Project partners are GFZ German Research Centre for Geosciences, Humboldt-Universität zu Berlin and Karlsruhe Institute of Technology (KIT). These three partners, with their expertise in information infrastructures, guarantee the sustainability of the registry.

Further information on re3data.org can be found in a recently published article in PLOS ONE:

Pampel, H., et al. (2013). Making Research Data Repositories Visible: The re3data.org Registry. PLOS ONE. doi: 10.1371/journal.pone.0078080


ORIGINAL: PLOS
By Heinz Pampel
November 4, 2013 

miércoles, 22 de mayo de 2013

Synthetic Biology Open Language Visual (SBOL Visual), version 1.0.0 RFC

ORIGINAL: MIT / SBOL

Synthetic Biology Open Language Visual (SBOL Visual), version 1.0.0
Download
Author: Quinn, Jacqueline; Beal, Jacob; Bhatia, Swapnil; Cai, Patrick; Chen, Joanna; Clancy, Kevin; Hillson, Nathan; Galdzicki, Michal; Maheshwari, Akshay; P, Umesh; Pocock, Matthew; Rodriguez, Cesar; Stan, Guy-Bart; Endy, Drew
Citable URI: http://hdl.handle.net/1721.1/78249
Date Issued: 2013-03-31
Abstract:
In this BioBricks Foundation Request for Comments (BBF RFC), we specify the Synthetic Biology Open Language Visual standard (SBOL Visual) to enable consistent, human-readable depiction of genetic designs.
URI: http://hdl.handle.net/1721.1/78249
Series/Report no.: BBF RFC;93
Keywords: exchange, data, biobrick

SBOL Visual Examples


ORIGINAL: SBOL
Synthetic Biology Open Language (SBOL) is an open-source standard for in silico representation of genetic designs. SBOL is designed to:
  • Allow synthetic biologists and genetic engineers to electronically exchange designs
  • Send and receive genetic designs to and from biofabrication centers
  • Facilitate storage of genetic designs in repositories
  • Embed genetic designs in publications
SBOL is built around the idea of a core that is used to unambiguously specify the design of a DNA molecule. Around the core are extensions that are used to increase the kind and amount of information transmitted by the language. There are six extensions under development. For example, one extension includes data and information on the performance of DNA components.

The adoption of SBOL offers many benefits, including: 
  1. enabling the use of multiple tools without rewriting designs for each tool, 
  2. enabling designs to be shared and published in a form other researchers can use even in a different software environment, and 
  3. ensuring the survival of design (and the intellectual effort put into them) beyond the lifetime of the software or the reseachers that were used to create them.
Adopting SBOL
SBOL comprises of an object model and a serialization of SBOL to a file. The file is a machine-readable format form representing designs in synthetic biology. SBOL is neutral with respect to programming languages and software encoding. By supporting SBOL for reading and writing synthetic biology designs, different software tools can directly communicate and store the same representation of these designs. This removes an impediment to sharing engineered systems and permits other researchers and commercial enterprises to start with an unambiguous representation of the design.

The supported serialization is SBOL:Core:rdf:xml. This SBOL serialization is a supported import and/or export format in many synthetic biology tools. GenBank files may be serialized in SBOL and SBOL serializations may be converted to Genbank files using JBEI's j5 SBOL XML <--> GenBank Conversion Utility.

If you are a software developer for synthetic biology, you should grab one of the libSBOL libraries. Java library libSBOLj can import and export the SBOL core file format v1.1. There is also a C/C++ based library, libSBOLc, being written by Jeffrey Johnson that can export and import the files in the same SBOL format. Therefore, Java developers can use the native Java library, and others can use the C/C++ library, either natively or by using bindings for other languages.

The examples page illustrates a variety of DNA designs specified using the core.
SBOL Developers
SBOL's development started in 2008 with a small grant from MS. Since then it has grown to include a wide consortium of individuals, public institutions, and commercial enterprises both in the US and Europe.

The SBOL Developers Group meets roughly twice a year to discuss progress of the standard. Work on libSBOL and SBOL's icrosoftvarious extensions is ongoing. To join the developers group, contact the Editors at 
sbol-editors@googlegroups.org.

miércoles, 23 de enero de 2013

Richard Gordon On Building Intelligent Machines

ORIGINAL: 33rdSquare
January 21, 2013

 

Richard T. Gordon is a specialist in machine learning, data analysis and the limits of predictability of systems. In a recent TEDx Talk, Gordon explains how artificial intelligence can be used to build and improve security systems. 

Ensuaratec's Richard T. Gordon has over 25 years of business experience, having held executive positions at NeuralSafe, Predictive Systems, Risk Management Solutions and the Chubb Corporation, as well as giving back to his community by spending two years working at environmental non-profit Bridging The Gap. Gordon has a PhD in physics and has spent his career working on complex problems affecting business, the nation’s computer infrastructure and the environment. He is a specialist in machine learning, data analysis and the limits of predictability of systems.

Gordon has published over 20 peer reviewed research papers including work on cyber security, intelligent computer systems, climate change and dynamic processes in the human brain. One of the principals in developing the models used by the financial industry to determine risk and exposure from catastrophes, his latest work is focused on the development of intelligent computer systemsto protect critical computer infrastructure from the complex cyber attacks that can be launched by rogue nations, terrorist groups and organized crime.

Gordon has a PhD in physics and has spent his career working on complex problems affecting business, the environment and the nation's computer infrastructure. Richard has been quoted or profiled Scientific American, Fortune Magazine, The Journal of Commerce and The Scientist.

In the TEDx Talk below, Gordon explains how artificial intelligence can be used to build and improve security systems.

SOURCE TEDx Talks

viernes, 16 de marzo de 2012

The Genomic Standards Consortium

ORIGINAL: PLoS Biology



Dawn Field1*, Linda Amaral-Zettler2, Guy Cochrane3,James R. Cole4, Peter Dawyndt5, George M. Garrity6, Jack Gilbert7,8, Frank Oliver Glöckner9, Lynette Hirschman10,Ilene Karsch-Mizrachi11, Hans-Peter Klenk12, Rob Knight13,Renzo Kottmann9, Nikos Kyrpides14, Folker Meyer7,15,Inigo San Gil16, Susanna-Assunta Sansone17, Lynn M. Schriml18, Peter Sterk19, Tatiana Tatusova11, David W. Ussery20, Owen White18, John Wooley21

  1. Centre for Ecology & Hydrology, Maclean Building, Crowmarsh Gifford, Wallingford, Oxfordshire, United Kingdom, 
  2. The Josephine Bay Paul Center for Comparative Molecular Biology and Evolution, Marine Biological Laboratory, Woods Hole, Massachusetts, United States of America, 
  3. European Molecular Biology Laboratory (EMBL) Outstation, European Bioinformatics Institute (EBI), Wellcome Trust Genome Campus, Hinxton, Cambridge, United Kingdom, 
  4. Center for Microbial Ecology, Michigan State University, East Lansing, Michigan, United States of America, 
  5. Department of Applied Mathematics and Computer Science, Ghent University, Ghent, Belgium, 
  6. Department of Microbiology and Molecular Genetics, Michigan State University, East Lansing, Michigan, United States of America, 
  7. Argonne National Laboratory, Argonne, Illinois, United States of America, 
  8. Department of Ecology and Evolution, University of Chicago, Chicago, Illinois, United States of America, 
  9. Microbial Genomics Group, Max Planck Institute for Marine Microbiology and Jacobs University Bremen, Bremen, Germany, 
  10. Information Technology Center, The MITRE Corporation, Bedford, Massachusetts, United States of America, 
  11. National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, United States of America, 
  12. DSMZ - German Collection of Microorganisms and Cell Cultures GmbH, Braunschweig, Germany, 
  13. Department of Chemistry and Biochemistry, University of Colorado, Boulder, Colorado, United States of America, 
  14. DOE Joint Genome Institute, Walnut Creek, California, United States of America, 
  15. Computation Institute, University of Chicago, Chicago, Illinois, United States of America, 
  16. LTER Network Office, Department of Biology, University of New Mexico, Albuquerque, New Mexico, United States of America, 
  17. University of Oxford, Oxford e-Research Centre, Oxford, United Kingdom, 
  18. Institute for Genome Sciences, University of Maryland School of Medicine, Baltimore, Maryland, United States of America, 
  19. Wellcome Trust Sanger Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge, United Kingdom, 
  20. Center for Biological Sequence Analysis, The Technical University of Denmark, Lyngby, Denmark, 
  21. University of California San Diego, La Jolla, California, United States of America
Abstract

A vast and rich body of information has grown up as a result of the world's enthusiasm for 'omics technologies. Finding ways to describe and make available this information that maximise its usefulness has become a major effort across the 'omics world. At the heart of this effort is the Genomic Standards Consortium (GSC), an open-membership organization that drives community-based standardization activities, Here we provide a short history of the GSC, provide an overview of its range of current activities, and make a call for the scientific community to join forces to improve the quality and quantity of contextual information about our public collections of genomes, metagenomes, and marker gene sequences.

Citation: Field D, Amaral-Zettler L, Cochrane G, Cole JR, Dawyndt P, et al. (2011) The Genomic Standards Consortium. PLoS Biol 9(6): e1001088. doi:10.1371/journal.pbio.1001088

Published: June 21, 2011

Copyright: © 2011 Field et al. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Funding: NERC International Opportunities Fund Award NE/3521773/1 and NE/E007325/1 (http://www.nerc.ac.uk/funding/) and National Science Foundation grant RCN4GSC, DBI-0840989 (http://www.nsf.gov/funding/). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.

Abbreviations: GSC, Genomic Standards Consortium; MIxS, Minimum Information about any (x) Sequence

* E-mail: dfield@ceh.ac.uk

Introduction

We currently have thousands of genomes, hundreds of metagenomes, and tens of thousands of marker gene data sets in the public domain, and these numbers are rapidly increasing [1]. Next-generation sequencing technologies promise to further fill the public databases with a bounty of information unthinkable even a few years ago. Each data set represents an organism or community with a unique biological history, sampling location, environmental context, and set of biologically interesting traits. Hence, each of these data sets makes a unique contribution to the ongoing creation of our public online catalogue of life.

We are now witnessing the rapid democratization of access to sequencing capacity—an immense opportunity for the global community, if proper stewardship of these data keeps pace [2],[3]. This stewardship must include enriching public sequence databases with the biological context of these sequences (Box 1), which will in turn necessitate the adoption of a fresh attitude to reporting results, both in our papers and our submissions to the public databases. Large, well-contextualized genome, metagenome, and marker gene data sets (e.g., ribosomal gene surveys) provide ideal opportunities for comparison and contrasting using computational means to solve a wide range of questions in biology (including questions in medicine, physiology, developmental biology, biogeochemistry, evolution, ecology, etc.).
Box 1. When the Cost of a Bacterial Genome Sequence Is Almost Nothing, That Organism's Contextual Information Is Increasingly Valuable

Consider the scenario where a new E. coli sequence has been obtained from a futuristic handheld device (like a Star Trek tricorder) that generates the complete genome in seconds. While the genome sequence may only be slightly different from strains already in the public databases, the metadata associated with this bug is both unique and crucial. Where and when was the E. coliisolated? Was it transmitted as a food-borne pathogen? Did it hospitalize the patient from whom it was isolated? Was it part of a larger infectious outbreak? Knowledge that a pathogen was isolated from diseased patients or healthy controls will readily assist in intervention strategies derived from machine-readable data.
These data sets should be treated as part of a larger whole—a catalogue of life on earth—that will allow us to observe, as we sample in time and space, how life changes. A range of ongoing and proposed megasequencing projects also promise to make great inroads into this grand vision (i.e.,

How must we now change the way we think about these data sets to prepare to integrate and co-analyze these large suites of related and contrasting data? Clearly, these data must be stored in robust comprehensive electronic systems that link to specific environments, diseases, or physiological states such that these relationships are electronically retrievable. To achieve this goal we urgently need shared standards that are both easy to use and scientifically robust.

The Genomic Standards Consortium
The GSC was established in late 2005 [9],[10] to tackle the challenge of working towards better descriptions of genomes and metagenomes through community-level, consensus-driven solutions. The GSC's mission is to work towards 1) the implementation of new genomic standards, 2) methods of capturing and exchanging the information captured in these standards (metadata, or contextual data) and 3) harmonization of information collection and analysis efforts across the wider genomics community.

The GSC fulfils this mission by holding face-to-face meetings, forming working groups, and building consensus products that can be widely used in this community. Thus far, the GSC has created a standard, the Minimum Information about any (x) Sequence (MIxS), that includes three minimum information checklists for describing genomes, metagenomes, and environmental marker sequences (MIGS/MIMS/MIMARKS) upon submission to the public databases and publication [11],[12]. MIxS requires core information on habitat, geolocation, and sequencing methodology as well as fields specific to data type and a range of optional environmental packages to capture core measurements defining a broad range of habitats, including water, soil, and host-associated habitats. The International Nucleotide Sequence Database Collaboration (INSDC; DDBJ/EMBL/GenBank) has created a GSC “keyword” (MIxS) to mark the richer entries complying with this standard.

Other working groups are dedicated to
  1. the maintenance of an extensible markup language (GCDML) that provides a reference implementation of the MIxS checklists [13]
  2. development of tools and software, 
  3. compliance and curation, and 
  4. biodiversity. 
Those requiring help complying with MIxS (curation support) should contact the compliance working group, and those requiring technical assistance in implementing/adopting these standards in software or database projects should contact the developer's working group (technical support). The developer and compliance groups work closely together, for example, to support compliance through a range of portals, including GOLD [1], MG-Rast[14], CAMERA [15], IMG/m [16], the RDP [17], SILVA [18], megx.net [19], and the ISA software suite[20]. The Biodiversity group works with communities to make sure that GSC standards evolve in harmony with standards for describing taxonomy and biodiversity.

The GSC has also stepped forward to create a journal designed to underpin the emerging field of standards development in the biological sciences [21]. The Standards in Genomic Sciences journal now serves as a formal voice for the GSC and supports the publication of standardized genome, metagenome, and pan-genome reports and other standards-supportive publications like Standard Operating Procedures (SOPs) [22] from the scientific community at large.

The GSC is now maturing into a hub for the coordination of large-scale projects. Two projects running under the GSC umbrella are the Microbial Earth Project, which calls for the coordinated sequencing of over 9,000 type strains (http://genome.jgi-psf.org/programs/bacte​ria-archaea/MEP/index.jsf), and the M5 project, which calls for the coordinated development of a next-generation computational infrastructure (http://gensc.org/gc_wiki/index.php/M5) [23].

The GSC also works closely with a range of related communities and helped drive the formation of the Environment Ontology [24], the Minimum Information for Biological and Biomedical Investigations (MIBBI) initiative [2], and most recently the BioSharing forum [3, 25].

A Call for Participation and Adoption

The Internet has resulted in a Cambrian explosion of productivity and data sharing through the adoption of a huge stack of agreed-upon protocols (standards) that allow many devices and programs to communicate to the transformative benefit of the everyday user [26]. Enabling access to user-generated content is key to harnessing the resources of a distributed community: Flickr has over 5 billion photographs uploaded, and Wikipedia has over 3.5 million English articles as of this writing. Standards for organizing sequence data will be similarly needed as sequencing instruments themselves, especially as these instruments are more and more commoditized and owned by individuals rather than institutions.

The tagline of the GSC is “Innovation through Collaboration”. For any standard to create a lasting impact requires substantial input from the wider scientific community, including adoption and support. The GSC urges researchers interested in pushing the boundaries of genomic science through collaboration to join and contribute expertise to building the GSC roadmap for the future. Membership in the GSC and all working groups is currently defined by participation. The GSC has a Board and several standing committees in addition to its working groups. For more information on the GSC, please see http://gensc.org/.

Conclusions
The GSC is working to become the authoritative working body in the area of genomics for the development and adoption of standards. We anticipate that the need for a collaborative body in which to build consensus at the community level and undertake large-scale projects will only increase with time, as in many ways the era of genomics is just beginning. In the future, sequence generation will only increase as access is further democratized. On one extreme, it will be like any other industrial commodity and will be outsourced into a global manufacturing marketplace. On the other, mid- to large-scale sequencing will be as locally accessible as a benchtop microscope or PCR machine is to a typical university researcher. Making these diverse streams of data accessible in a coherent framework will require new, standardized ways of describing, storing, and exchanging this information. The framework required to do this will involve acceptance of profound sociological and technological changes in how we do business in the genomic sciences.

Acknowledgments
The GSC acknowledges all participants in past GSC meetings for their thoughtful contributions. The GSC also acknowledges a range of funding sources for its past meetings, including NERC, NIEeS, NSF, the Gordon and Betty Moore Foundation and DOE. In particular, funding from NERC helped launch the GSC and allow essential infrastructure to be built. Funding from the NSF in the form of the Research Co-ordination Network (RCN4GSC) is supporting exchange visits of early career scientists and working group activities.

References
  1. Liolios K, Chen I. M, Mavromatis K, Tavernarakis N, Hugenholtz P, et al. (2010) The Genomes On Line Database (GOLD) in 2009: status of genomic and metagenomic projects and their associated metadata. Nucleic Acids Res 38: D346–D354. FIND THIS ARTICLE ONLINE
  2. Taylor C. F, Field D, Sansone S. A, Aerts J, Apweiler R, et al. (2008) Promoting coherent minimum reporting guidelines for biological and biomedical investigations: the MIBBI project. Nat Biotechnol 26: 889–896. FIND THIS ARTICLE ONLINE
  3. Field D, Sansone S. A, Collis A, Booth T, Dukes P, et al. (2009) 'Omics data sharing. Science 326: 234–236. FIND THIS ARTICLE ONLINE
  4. Wu D, Hugenholtz P, Mavromatis K, Pukall R, Dalin E, et al. (2009) A phylogeny-driven genomic encyclopaedia of Bacteria and Archaea. Nature 462: 1056–1060. FIND THIS ARTICLE ONLINE
  5. Turnbaugh P. J, Ley R. E, Hamady M, Fraser-Liggett C. M, Knight R, et al. (2007) The human microbiome project. Nature 449: 804–810. FIND THIS ARTICLE ONLINE
  6. Gilbert J, Meyer F, Jansson J, Gordon J, Pace N, et al. (2010) The Earth Microbiome Project: Meeting report of the “1st EMP meeting on sample selection and acquisition” at Argonne National Laboratory October 6th 2010. Stand Genomic Sci 3: 249–253. FIND THIS ARTICLE ONLINE
  7. (2009) Genome 10K: a proposal to obtain whole-genome sequence for 10,000 vertebrate species. J Hered 100: 659–674. FIND THIS ARTICLE ONLINE
  8. Rusch D. B, Halpern A. L, Sutton G, Heidelberg K. B, Williamson S, et al. (2007) The Sorcerer II Global Ocean Sampling expedition: northwest Atlantic through eastern tropical Pacific. PLoS Biol 5: e77. doi:10.1371/journal.pbio.0050077.
  9. Field D, Hughes J (2005) Cataloguing our current genome collection. Microbiology 151: 1016–1019. FIND THIS ARTICLE ONLINE
  10. Field D, Garrity G, Morrison N, Selengut J, Sterk P, et al. (2005) eGenomics: cataloguing our complete genome collection. Comp Funct Genomics 6: 363–368. FIND THIS ARTICLE ONLINE
  11. Field D, Garrity G, Gray T, Morrison N, Selengut J, et al. (2008) The minimum information about a genome sequence (MIGS) specification. Nat Biotechnol 26: 541–547. FIND THIS ARTICLE ONLINE
  12. Yilmaz P, Kottmann R, Field D, Knight R, Cole J. R, et al. (2011) The “Minimum Information about a MARKer gene Sequence” (MIMARKS) specification. Nat Biotechnol 29: 415–420. FIND THIS ARTICLE ONLINE
  13. Kottmann R, Gray T, Murphy S, Kagan L, Kravitz S, et al. (2008) A standard MIGS/MIMS compliant XML Schema: toward the development of the Genomic Contextual Data Markup Language (GCDML). OMICS 12: 115–121. FIND THIS ARTICLE ONLINE
  14. Meyer F, Paarmann D, D'Souza M, Olson R, Glass E. M, et al. (2008) The metagenomics RAST server - a public resource for the automatic phylogenetic and functional analysis of metagenomes. BMC Bioinformatics 9: 386. FIND THIS ARTICLE ONLINE
  15. Sun S, Chen J, Li W, Altinatas I, Lin A, et al. (2011) Community cyberinfrastructure for Advanced Microbial Ecology Research and Analysis: the CAMERA resource. Nucleic Acids Res 39: D546–D551.FIND THIS ARTICLE ONLINE
  16. Markowitz V. M, Chen I. M, Palaniappan K, Chu K, Szeto E, et al. (2010) The integrated microbial genomes system: an expanding comparative analysis resource. Nucleic Acids Res 38: D382–D390.FIND THIS ARTICLE ONLINE
  17. Cole J. R, Wang Q, Cardenas E, Fish J, Chai B, et al. (2009) The Ribosomal Database Project: improved alignments and new tools for rRNA analysis. Nucleic Acids Res 37: D141–D145. FIND THIS ARTICLE ONLINE
  18. Pruesse E, Quast C, Knittel K, Fuchs B. M, Ludwig W, et al. (2007) SILVA: a comprehensive online resource for quality checked and aligned ribosomal RNA sequence data compatible with ARB. Nucleic Acids Res 35: 7188–7196. FIND THIS ARTICLE ONLINE
  19. Kottmann R, Kostadinov I, Duhaime M. B, Buttigieg P. L, Yilmaz P, et al. (2010) Megx.net: integrated database resource for marine ecological genomics. Nucleic Acids Res 38: D391–395.FIND THIS ARTICLE ONLINE
  20. Rocca-Serra P, Brandizi M, Maguire E, Sklyar N, Taylor C, et al. (2010) ISA software suite: supporting standards-compliant experimental annotation and enabling curation at the community level. Bioinformatics 26: 2354–2356. FIND THIS ARTICLE ONLINE
  21. Garrity G. M, Field D, Kyrpides N, Hirschman L, Sansone S. A, et al. (2008) Toward a standards-compliant genomic and metagenomic publication record. OMICS 12: 157–160. FIND THIS ARTICLE ONLINE
  22. Angiuoli S. V, Gussman A, Klimke W, Cochrane G, Field D, et al. (2008) Toward an online repository of Standard Operating Procedures (SOPs) for (meta)genomic annotation. OMICS 12: 137–141.FIND THIS ARTICLE ONLINE
  23. (2009) Metagenomics versus Moore's law. Nat Meth 6: 623. FIND THIS ARTICLE ONLINE
  24. Morrison N, Wood A. J, Hancock D, Shah S, Hakes L, et al. (2006) Annotation of environmental OMICS data: application to the transcriptomics domain. OMICS 10: 172–178. FIND THIS ARTICLE ONLINE
  25. Field D, Sansone S, Delong E. F, Sterk P, Friedberg I, et al. (2010) Meeting report: BioSharing at ISMB 2010. Stand Genomic Sci 3: 254–258. FIND THIS ARTICLE ONLINE
  26. Berners-Lee T (22 November 2010) Long live the web: a call for continued open standards and neutrality. Scientific American. FIND THIS ARTICLE ONLINE