Mostrando entradas con la etiqueta Analysis. Mostrar todas las entradas
Mostrando entradas con la etiqueta Analysis. Mostrar todas las entradas

viernes, 16 de octubre de 2015

Automating big-data analysis. System that replaces human intuition with algorithms outperforms 615 of 906 human teams.


COMMENT
Big-data analysis consists of searching for buried patterns that have some kind of predictive power. But choosing which “features” of the data to analyze usually requires some human intuition. In a database containing, say, the beginning and end dates of various sales promotions and weekly profits, the crucial data may not be the dates themselves but the spans between them, or not the total profits but the averages across those spans.

MIT researchers aim to take the human element out of big-data analysis, with a new system that not only searches for patterns but designs the feature set, too. To test the first prototype of their system, they enrolled it in three data science competitions, in which it competed against human teams to find predictive patterns in unfamiliar data sets. Of the 906 teams participating in the three competitions, the researchers’ “Data Science Machine” finished ahead of 615.

In two of the three competitions, the predictions made by the Data Science Machine were 94 percent and 96 percent as accurate as the winning submissions. In the third, the figure was a more modest 87 percent. But where the teams of humans typically labored over their prediction algorithms for months, the Data Science Machine took somewhere between two and 12 hours to produce each of its entries.

We view the Data Science Machine as a natural complement to human intelligence,” says Max Kanter, whose MIT master’s thesis in computer science is the basis of the Data Science Machine. “There’s so much data out there to be analyzed. And right now it’s just sitting there not doing anything. So maybe we can come up with a solution that will at least get us started on it, at least get us moving.

Between the lines
Kanter and his thesis advisor, Kalyan Veeramachaneni, a research scientist at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), describe the Data Science Machine in a paper that Kanter will present next week at the IEEE International Conference on Data Science and Advanced Analytics.

Veeramachaneni co-leads the Anyscale Learning for All group at CSAIL, which applies machine-learning techniques to practical problems in big-data analysis, such as determining the power-generation capacity of wind-farm sites or predicting which students are at risk for dropping out of online courses.

What we observed from our experience solving a number of data science problems for industry is that one of the very critical steps is called feature engineering,” Veeramachaneni says. “The first thing you have to do is identify what variables to extract from the database or compose, and for that, you have to come up with a lot of ideas.

In predicting dropout, for instance, two crucial indicators proved to be how long before a deadline a student begins working on a problem set and how much time the student spends on the course website relative to his or her classmates. MIT’s online-learning platform MITx doesn’t record either of those statistics, but it does collect data from which they can be inferred.

Featured composition
Kanter and Veeramachaneni use a couple of tricks to manufacture candidate features for data analyses.

  • One is to exploit structural relationships inherent in database design. Databases typically store different types of data in different tables, indicating the correlations between them using numerical identifiers. The Data Science Machine tracks these correlations, using them as a cue to feature constructionFor instance, one table might list retail items and their costs; another might list items included in individual customers’ purchases. The Data Science Machine would begin by importing costs from the first table into the second. Then, taking its cue from the association of several different items in the second table with the same purchase number, it would execute a suite of operations to generate candidate features
    • total cost per order, 
    • average cost per order, 
    • minimum cost per order, and 
    • so on. As numerical identifiers proliferate across tables, the Data Science Machine layers operations on top of each other, finding minima of averages, averages of sums, and so on.
  • It also looks for so-called categorical data, which appear to be restricted to a limited range of values, such as days of the week or brand names. It then generates further feature candidates by dividing up existing features across categories. Once it’s produced an array of candidates, it reduces their number by identifying those whose values seem to be correlated. Then it starts testing its reduced set of features on sample data, recombining them in different ways to optimize the accuracy of the predictions they yield.
The Data Science Machine is one of those unbelievable projects where applying cutting-edge research to solve practical problems opens an entirely new way of looking at the problem,” says Margo Seltzer, a professor of computer science at Harvard University who was not involved in the work. “I think what they’ve done is going to become the standard quickly — very quickly.

ORIGINAL: MIT News
Larry Hardesty | MIT News Office 
October 16, 2015

lunes, 10 de noviembre de 2014

Rosetta's comet lander readies for its launch

Rosetta that was launched in 2004 is now preparing to enter its last phase of the mission. Ten years later, on August 6, the spacecraft began orbiting 67P, and its 11 instruments started scrutinizing myriad characteristics of the comet (SN: 9/6/14, p. 8). Those instruments, plus the cameras and sensors on the Philae lander, are designed to map 67P, determine what it’s made of and observe how its chemistry might change as it swings around the sun.

On November 12, Rosetta will sidle up to a comet, steady itself and drop a 100-kilogram robotic lander toward the hunk of rock, dust and ice. The lander, named Philae, will drift through space, tugged only slightly by the gravity of the comet, commonly called 67P. Mission scientists will be holding their breath for what could be several anxiety-filled hours to see if Philae lands where and how it’s supposed to.

The exercise — the first attempt to set a lander on a comet — is as nerve-racking as landing on Mars or the moon, with some added challenges. Comets and other small space rocks have much less gravity than planets or moons, which is why it will take Philae close to seven hours to float to comet 67P’s surface. Then there’s the comet’s speed: Rosetta will drop the lander toward 67P as the comet shoots through the solar system at 55,000 kilometers per hour.

Add to that a comet's unpredictable nature: At any moment and without warning, 67P might spew out jets of gas and dust. Such eruptions could blow the spacecraft off course or skew the lander’s trajectory so it hits a boulder or misses its mark.

Early in the mission, scientists estimated that Philae had a 70 to 75 percent chance of successfully touching down on the comet, officially known as 67P/Churyumov-Gerasimenko. They made that prediction when they thought the comet was shaped like a potato. In July, Rosetta began sending pictures of 67P, indicating it looks more like a rubber duck — two masses connected by a thin neck. The new shape adds a bit more uncertainty to Philae sticking its landing.

Video Transcript :
"The European Space Agency's Philae robot is preparing for an extraordinary landing on an ambitious target: comet 67P/Churyumov-Gerasimenko.

Out there in the cold depths of space, somewhere between Mars and Jupiter, a spacecraft called Rosetta is coyly playing with a comet. Scientists call this comet 67P/Churyumov-Gerasimenko. Since August, the spacecraft has been swooping in and around the 4-kilometer long comet, snapping selfies with it and taking close-ups of its craggy cliffs, crevices and boulders.

These striking images are the first to clearly show what the surface of comets can look like. They also helped scientists pinpoint just the right spot for a daring operation: To set a lander down on 67P’s surface.

After weeks of pouring over the images mission scientists finally selected a relatively flat spot on the head of the comet for their lander, Philaeto settle into. But getting down to the surface won’t be easy.

It will require a complex set of circles around the comet and a separation at just the right point in one of those orbits.

This separation is slated to take place on November 12. That day, Rosetta will swing to within 23 kilometers of the comet and drop Philae off its backside. The lander will then drift through space toward 67P. And, if all goes well, it will arrive on the comet roughly seven hours later.

Then the 10 instruments aboard the robot will probe the comet from the inside out. Each instrument is designed to look at specific features of the comet, such as its internal structure and the chemical elements and molecules that make up its dust. This data, along with Rosetta’s analysis of 67P, will help scientists really get a handle on what comets are and what happens to them as their orbits take them closer to the sun.

A close look at the comet could also take us back in time to the beginning of the solar system. That’s because comets appear to be time capsules that may have preserved some of the earliest materials found in the solar system. Looking at those pristine features may help scientists and astronomers answer fundamental questions about how the planets formed and possibly even how water and other life ingredients made their way to Earth.

But first we’ll have to wait for that long-sought signal saying yes, Philae has made its touchdown.


Images, graphics and animations courtesy of DLR German Aerospace Center and ESA; Narrated and produced by Ashley Yeager"





ORIGINAL: Science Dump
by Andreea
11/08/2014