Mostrando entradas con la etiqueta Deep Learning. Mostrar todas las entradas
Mostrando entradas con la etiqueta Deep Learning. Mostrar todas las entradas

domingo, 26 de agosto de 2018

Millimeter-Scale Computers: Now With Deep Learning Neural Networks on Board

Photo: University of Michigan and TSMCOne of several varieties of University of Michigan micro motes. This one incorporates 1 megabyte of flash memory.
Computer scientist David Blaauw pulls a small plastic box from his bag. He carefully uses his fingernail to pick up the tiny black speck inside and place it on the hotel café table. At one cubic millimeter, this is one of a line of the world’s smallest computers. I had to be careful not to cough or sneeze lest it blow away and be swept into the trash.

Blaauw and his colleague Dennis Sylvester, both IEEE Fellows and computer scientists at the University of Michigan, were in San Francisco this week to present ten papers related to these “micro mote” computers at the IEEE International Solid-State Circuits Conference (ISSCC). They’ve been presenting different variations on the tiny devices for a few years.

Their broader goal is to make smarter, smaller sensors for medical devices and the internet of things—sensors that can do more with less energy. Many of the microphones, cameras, and other sensors that make up eyes and ears of smart devices are always on alert, and frequently beam personal data into the cloud because they can’t analyze it themselves. Some have predicted that by 2035, there will be 1 trillion such devices. “If you’ve got a trillion devices producing readings constantly, we’re going to drown in data,” says Blaauw. By developing tiny, energy efficient computing sensors that can do analysis on board, Blaauw and Sylvester hope to make these devices more secure, while also saving energy.


Photo: University of Michigan/TSMCMade of multiple layers of computing.

At the conference, they described micro mote designs that use only a few nanowatts of power to perform tasks such as distinguish the sound of a passing car and measuring temperature and light levels. They showed off a compact radio that can send data from the small computers to receivers 20 meters away—a considerable boost compared to the 50 centimeter range they reported last year at ISSCC. They also described their work with TSMC on embedding flash memory into the devices, and a project to bring on board dedicated, low-power hardware for running artificial intelligence algorithms called deep neural networks.

Blaauw and Sylvester say they take a holistic approach to adding new features without ramping up power consumption. “There’s no one answer” to how the group does it, says Sylvester. If anything, it’s “smart circuit design,” Blaauw adds. (They pass ideas back and forth rapidly, not finishing each other’s sentences but something close to it.)

The memory research is a good example of how the right tradeoffs can improve performance, says Sylvester. Previous versions of the micro motes used 8 kilobytes of SRAM, which makes for a pretty low-performance computer. To record video and sound, the tiny computers need more memory. So the group worked with TSMC to bring flash memory on board. Now they can make tiny computers with 1 megabyte of storage.



Flash can store more data in a smaller footprint than SRAM, but it takes a big burst of power to write to the memory. With TSMC, the group designed a new memory array that uses a more efficient charge pump for the writing process. The memory arrays are a bit less dense than TSMC’s commercial products, for example, but still much better than SRAM. “We were able to get huge gains with small trade-offs,” says Sylvester.

Another micro mote they presented at the ISSCC incorporates a deep-learning processor that can operate a neural network while using just 288 microwatts. Neural networks are artificial intelligence algorithms that perform well at tasks such as face and voice recognition. They typically demand both large memory banks and intense processing power, and so they’re usually run on banks of servers often powered by advanced GPUs. Some researchers have been trying to lessen the size and power demands of deep-learning AI with dedicated hardware that’s specially designed to run these algorithms. But even those processors still use over 50 milliwatts of power—far too much for a micro mote. The Michigan group brought down the power requirements by redesigning the chip architecture, for example by situating four processing elements within the memory (in this case, SRAM) to minimize data movement.

The idea is to bring neural networks to the internet of things. “A lot of motion detection cameras take pictures of branches moving in the wind—that’s not very helpful,” says Blaauw. Security cameras and other connected devices are not smart enough to tell the difference between a burglar and a tree, so they waste energy sending uninteresting footage to the cloud for analysis. On-board deep-learning processors could make better decisions, but only if they don’t use too much power. The Michigan group imagine deep-learning processors could be integrated into many other internet-connected things besides security systems. For example, an HVAC systems could decide to turn the air conditioning down if they see multiple people putting on their coats.

After demonstrating many variations on these micro motes in an academic setting, the Michigan group hopes they will be ready for market in a few years. Blaauw and Sylvester say their start-up company CubeWorks is currently prototyping devices and researching markets. The company was quietly incorporated in late 2013. Last October, Intel Capital announced they had invested an undisclosed amount in the tiny computer company. 




Posted 10 Feb 2017


lunes, 20 de agosto de 2018

You Should Know These 20 Technology Leaders Driving China's A.I. Revolution


China’s leading technology companies are on fire, heavily investing in artificial intelligence and building true global presences. McKinsey recently reported that academic and research institutions in the country publish more cited research papers than the US, UK, or any other global leader in AI, producing nearly 10,000 papers in 2015 alone.

Backed by strong government mandates and billions of dollars of both private and public investments, China is challenging the US for position of global AI leader. Fearful of competition, the US government is considering placing restrictions on Chinese investments in AI and technology in the United States. In many sectors, such as healthcare, China may already be ahead of America in applying AI to critical public issues.

You might recognize names like Andrew Ng, Sebastian Thrun, Geoffrey Hinton, or Yann LeCun as important figures in AI, but few Westerners can name the key leaders driving AI innovation in China and at Chinese companies globally. These executives, entrepreneurs, professors, and researchers helm the most important Chinese tech companies and research labs and are respected widely for their technical expertise and accomplishments.

We’ve researched and curated 20 of the most important figures in the Chinese AI landscape that you should know

1. KAI-FU LEE
Co-Founder of Sinovation Ventures, Former President of Google China

Kai-Fu Lee is a globally recognized technology leader with executive experience at Apple, Microsoft, and Google. He got his BS in Computer Science from Columbia University and his PhD from Carnegie Mellon. Lee established Google China prior to co-founding Sinovation Ventures, a venture capital firm actively funding technology and AI startups in the US and China.

With celebrity status in China and over 50 million followers on Chinese social networks, Lee has become an oracle in predicting trends in Chinese tech. Lee told CNBC recently that artificial intelligence is the “singular thing that will be larger than all of human tech revolutions added together, including electricity, the industrial revolution, internet, and mobile internet.

2. QI LU
Group President & COO, Baidu


Qi Lu was hired by Baidu to lead the company’s strategic efforts in AI and push forward integration and collaboration within the company. Every Baidu business unit, including AI teams working on autonomous driving, reports to Lu. A spokesperson from Baidu stated: “With Dr. Lu on board, we are confident that our strategy will be executed smoothly and Baidu will become a world-class technology company and global leader in AI.

Prior to joining Baidu, Lu was personally recruited by Steve Ballmer to join Microsoft where he eventually became EVP of the Applications & Services Group. Lu started his professional career in IBM’s research labs, before joining Yahoo and rising to EVP of the Search & Advertising Group. He completed a BS in Computer Science at Fudan University and was invited by Carnegie Mellon professor Edmund M. Clarke to pursue his PhD at CMU.


3. HAIFENG WANG
Head of AI Group, Baidu

After Andrew Ng’s departure from Baidu, Haifeng Wang took over as leader of the expanded AI Group (AIG), consisting of
  • Baidu’s Institute of Deep Learning, 
  • Big Data Lab, 
  • Silicon Valley AI Lab, 
  • Augmented Reality Lab, 
  • Natural Language Unit, 
  • AI Platform Unit, and 
  • a few other departments.
Wang’s technical specialty is natural language processing (NLP) and machine translation and he has authored over 100 academic papers in AI. He applies his expertise to Baidu’s efforts in
  • NLP, 
  • computer vision, 
  • speech recognition, 
  • knowledge graphs, 
  • personalized recommendations, and 
  • deep learning. 
Wang is also an adjunct professor at Harbin Institute of Technology where he received his BS, MS, and PhD degrees in Computer Science.

4. TONG ZHANG
Executive Director of AI Lab, Tencent



The battle for top AI talent is incredibly fierce. Tong Zhang was poached from Baidu by Tencent last year to lead Tencent’s newly established AI lab. Formerly he was head of Baidu’s Big Data Lab, worked at IBM and Yahoo, and was a professor at Rutgers University.

With a team of over 200 engineers, Zhang is focused on developing Tencent’s capabilities in machine learning, computer vision, speech recognition, and natural language processing and applying new AI technologies to the company’s vast array of popular consumer products like WeChat.

5. JINGREN ZHOU
Chief Scientist and Vice President of Alibaba Cloud, Alibaba


Alibaba Cloud launched in 2009 and is now Alibaba’s fastest growing business unit. Similar to Amazon Web Services (AWS), Alibaba Cloud, also called Aliyun, emerged out of the company’s need for enormous computing power to handle millions of online shopping transactions.

Jingren Zhou leads big data and AI research at Alibaba Cloud’s Institute of Data Science Technology (iDST). In this role, he drives Alibaba’s AI technologies in speech, natural language, image and video processing, and large-scale machine learning.

Prior to joining Alibaba, Zhou was an engineering manager at Microsoft in charge of developing the big data computation platform supporting Windows, Office, and Bing. He received his BS from the University of Science and Technology of China and his PhD in Computer Science from Columbia University.

6. XIAOFE HE
President, DiDi Research


DiDi Chuxing is the “Uber of China”, with over 50TB of real-time data and over 9 billion routes driven per day. DiDi Research, the “brains of Didi Chuxing”, is a machine learning research institute set up by the company to predict demand, reduce surge impact, and also develop self-driving car technology.

President Xiaofe He got his BS in Computer Science from Zhejiang University and PhD from University of Chicago. Prior to helming DiDi Research, he worked as a research scientist and President of Yahoo Research Labs and joined Zhejiang University as a professor focused on applying mathematics and data analysis to solve important problems in pattern recognition, multimedia, and computer vision.

7. YUANQING LIN
Head of Baidu Research, Baidu


As Head of Baidu Research, Yuanqin Lin manages Baidu’s research labs, which include the

Along with Wei Xu, he will be leading Baidu’s contributions to China’s government-funded National Engineering Laboratory of Deep Learning Technology that is co-helmed by Tsinghua & Beihang University.

Prior to Baidu, Lin was the head of Media Analytics at NEC Labs America where he led teams focusing on computer vision research for mobile search and driverless cars. Lin received his MS degree in Optical Engineering from Tsinghua University and his PhD in Electrical Engineering from University of Pennsylvania.

8. PINPIN ZHU
President & CTO, Xiaoi


Xiaoi is China’s leading platform for conversational AI, powering the majority of the country’s bot and virtual assistant experiences. Established in Shanghai in 2001, the company’s technologies are used by hundreds of medium to large enterprises, government entities, and over 500 million users collectively.

Pinpin Zhu’s numerous patents in the space – including ones for “Chatting Robot System” and “SMS Robot System” – drove Xiaoi’s technical dominance in conversational interfaces. In addition to running Xiaoi, Zhu is also a Doctor of Science at the Chinese Academy of Sciences, has been appointed to China National Information Technology Standardization Committee, and has received numerous awards and accolades for his contributions to the field.

9. WEI XU
Distinguished Scientist, Baidu


In a company full of highly credentialed scientists, researchers, and engineers, Wei Xu is the only one with the title “Distinguished Scientist”. He is highly respected for his technical chops within the company due to his work on PaddlePaddle, a deep learning toolkit which was open sourced in late 2016. In development for over three years, PaddlePaddle is used to power search rankings, targeted advertising, image classification, translation, and self-driving cars.

Xu received his Bachelor’s degree at Tsinghua University, his MS from Carnegie Mellon, and was previously a researcher at NEC Labs and Facebook before joining Baidu.

10. WANLI MIN
Principal Data Scientist, Alibaba


Wanli Min led the research and development of Alibaba Cloud’s (Aliyun) artificial intelligence system, named Little Ai. Ai has been deployed by Alibaba internally to support customer service and traffic pattern predictions for the company’s flagship e-commerce business. Min also used machine learning to predict the winner of a top-rated Chinese reality TV show called “I Am Singer” and helped city planners in Guangdong province optimize traffic lights in real-time to reduce congestion.

Min entered college at the age of 14 and received his Bachelors from the University of Science & Technology of China and a PhD in Statistics from the University of Chicago.

11. KUN JING
General Manager of Duer, Baidu


Duer is Baidu’s answer to Apple’s Siri, Amazon’s Alexa, Microsoft’s Cortana, and Google’s Assistant. The conversational AI platform powers virtual assistant capabilities in a number of devices, ranging from XiaoYu, China’s version of the Amazon Echo, to voice-activated smart televisions.

Kun Jing leads the Duer business unit. Prior to joining Baidu, Jing was Microsoft’s R&D Director and created Xiaoice, a popular chatbot that went viral on Tencent’s WeChat and Sina’s Weibo. Xiaoice has over 20 million registered users who interact with the bot an average of 60 times a month, earning it the rank of Weibo’s top influencer.

12. DONG YU
Deputy Director of AI Lab, Tencent


Hired as deputy head of Tencent’s AI Lab, Dong Yu co-runs the new lab with Tong Zhang and spearheads research in speech recognition and natural language understanding. Prior to joining Tencent, Yu was the principal researcher at Microsoft Research Institute’s Speech and Dialog Group, an adjunct professor at Zhejiang University, a visiting professor at University of Science and Technology of China, and a visiting researcher at Shanghai Jiao Tong University. He received a Bachelor’s in Electrical Engineering from Zhejiang University and a PhD in Computer Science from Idaho University.

“I’m excited to join AI Lab,” Yu shares. “Over the past decade, Tencent has accumulated abundant experience in application scenarios, developed a massive data bank, established powerful computing capabilities, and built an outstanding team of technology experts; all which have helped form the foundation of in-depth research and AI application at Tencent today.”

13. ADAM COATES
Director of Silicon Valley AI Lab, Baidu


Coates received his BS, MS, and PhD degrees in Computer Science from Stanford University and has worked on everything from computer vision for autonomous cars, deep learning for speech recognition, and machine learning for helicopter acrobatics. At Baidu, he worked on DeepSpeech, a speech recognition and transcription engine that performs as well as native Mandarin speakers, and DeepVoice, a text-to-speech synthesis engine that generates believable human-like audio.

Coates is particularly excited about putting AI in the hands of real-world consumers. When he was selected by MIT Technology Review as one of 35 Innovators Under 35 in 2015, he explained that “in rapidly developing economies like in China, there are many people who will be connecting to the Internet for the first time through a mobile phone. Having a way to interact with a device or get the answer to a question as easily as talking to a person is even more powerful to them. I think of Baidu’s customers as having a greater need for artificial intelligence than myself.

14. KAI YU
Founder & CEO, Horizon Robotics


Formerly head of Baidu’s Institute of Deep Learning, Kai Yu left Baidu to start Beijing-based startup Horizon Robotics. Funded by leading investors like Yuri Milner and Sequoia Capital, Yu’s mission is to become the “Android of Robotics,” a pervasive AI system that powers all of our smart devices. Unlike other Chinese tech giants which dominate in the cloud, Horizon aims to adapt AI to every piece of hardware in the physical world.
Horizon has launched two platforms to date: 
  • Anderson for smart homes and 
  • Hugo for smart driving. 
Anderson imbues home appliances with capabilities such as facial recognition and automatic ordering, while Hugo is an advanced driver assistance system that performs real-time pedestrian and object detection even in adverse weather conditions.

Yu received his BS and MS degrees in Electrical Engineering from Nanjing University and his PhD in Computer Science from Ludwig-Maximilians Universitat Munchen in Germany.

15. JING WANG
Former Senior Vice President of Engineering, Baidu


While at Baidu, Jing Wang managed over 5,000 engineers in numerous business units, including the ones he founded:
  • Mobile, 
  • Cloud Computing, 
  • Big Data, 
  • Cybersecurity, 
  • Baidu Research, and 
  • Autonomous Driving. 
He left the company shortly after Andrew Ng’s resignation to start his own self-driving car company, and is widely credited with driving forward Baidu’s progress in the space.

Prior to joining Baidu, Wang was Deputy Head of Google’s Shanghai engineering office as well as eBay China’s CTO and R&D general manager. He received his Bachelor’s from the University of Science & Technology of China and his Master’s in Computer Science from the Chinese Academy of Sciences.

16. BO ZHANG
Professor of Computer Science and Technology, Tsinghua University


As a professor at Tsinghua University, Bo Zhang’s research interests include AI, machine learning, pattern recognition, knowledge engineering, and robotics. His notable academic achievements include advances in robotic task and motion planning, probabilistic logic neural networks (PLN), and machine learning algorithms for image retrieval and classification and webpage structure mining.

Along with Baidu and Wei Li of Beihang University, Tsinghua was selected to co-lead the government-funded National Engineering Laboratory of Deep Learning. He is a member of the Chinese Academy of Sciences and received his Bachelor’s in Automatic Control from Tsinghua University.

17. HUA WU
Technical Chief of NLP Group, Baidu


Hua Wu contributed a number of technical breakthroughs in 
  • natural language processing (NLP), 
  • dialogue systems, and 
  • neural machine translation (NMT) 21
in her seven year tenure at Baidu. The New York Times hailed her research work in multi-task learning as “pathbreaking” and she was able to successfully deploy her invention at scale to hundreds of millions of users of Baidu’s translation products. Wu is also responsible for the technology behind Baidu’s conversational AI, Duer.

Wu received her PhD from the Chinese Academy of Sciences and co-chairs leading academic AI conferences such as ACL and IJCAI.

18. WEI LI
President and Professor of Computer Science, Beihang University


Along with Bo Zhang of Tsinghua University and senior executives from Baidu, Wei Li was selected to co-lead China’s National Engineering Laboratory of Deep Learning. He is a member of the Chinese Academy of Sciences and also president of Beihang University. Li has won numerous accolades and prizes for his technical contributions in artificial intelligence and network computing.

Li graduated from the Department of Mathematics and Mechanics of Beijing University and received his PhD in Computer Science from the University of Edinburgh.

19. HONGBIN ZHA
Professor of Machine Learning, Peking University 


China hopes to leap-frog the US and other Western countries by vast and fast investment in the AI industry,says Hongbin Zha, AI researcher and professor at Beijing’s Peking University. Zha directs the Key Lab of Machine Perception at Peking University and collaborates with Microsoft Research Asia alongside other AI leaders across the continent. His research interests include computer vision theory, virtual reality, and robotics.

He received his Bachelor’s degree in Electrical Engineering from Hefei University of Technology in China and his MS and PhD degrees in Electrical Engineering from Kyushu University in Japan.

20. YUNJI CHEN


In 2015, Yunji Chen was selected by MIT Technology Review as one of their top 35 Innovators Under 35. Described as “iconoclastic and cosmopolitan”, he was chosen for his work in designing specialized deep-learning processors which dramatically reduce the computational costs of large-scale machine learning. His dream is to enable even common cell phones to be “as powerful as Google Brain”.

Chen entered college at age 14 and completed his PhD with lightning speed by the age of 24. He’s now chief architect of the Godson-3C, a microprocessing chip that reduces energy requirements for computers to recognize objects and translate languages and is developing the Cambricon, a brain-inspired processor chip that models human nerve cells and synapses to facilitate deep learning. The research team is led by Chen and his younger brother, Tianshi Chen, two of the youngest professors at the Chinese Academy of Sciences.


ABOUT THE AUTHOR
Adelyn is the Head of Marketing at TOPBOTS. She's got a decade of experience growing billion-dollar companies like Eventbrite, NextDoor, and Amazon. Follow her on Twitter at @adelynzhou to learn how to accelerate your growth with AI.

ORIGINAL: TopBots
Jun 18, 2017

miércoles, 1 de marzo de 2017

Google Unveils Neural Network with “Superhuman” Ability to Determine the Location of Almost Any Image

Guessing the location of a randomly chosen Street View image is hard, even for well-traveled humans. But Google’s latest artificial-intelligence machine manages it with relative ease.
Here’s a tricky task. Pick a photograph from the Web at random. Now try to work out where it was taken using only the image itself. If the image shows a famous building or landmark, such as the Eiffel Tower or Niagara Falls, the task is straightforward. But the job becomes significantly harder when the image lacks specific location cues or is taken indoors or shows a pet or food or some other detail.

Nevertheless, humans are surprisingly good at this task. To help, they bring to bear all kinds of knowledge about the world such as the type and language of signs on display, the types of vegetation, architectural styles, the direction of traffic, and so on. Humans spend a lifetime picking up these kinds of geolocation cues.

So it’s easy to think that machines would struggle with this task. And indeed, they have.

Today, that changes thanks to the work of Tobias Weyand, a computer vision specialist at Google, and a couple of pals. These guys have trained a deep-learning machine to work out the location of almost any photo using only the pixels it contains.

Their new machine significantly outperforms humans and can even use a clever trick to determine the location of indoor images and pictures of specific things such as pets, food, and so on that have no location cues.

Their approach is straightforward, at least in the world of machine learning. 
  • Weyand and co begin by dividing the world into a grid consisting of over 26,000 squares of varying size that depend on the number of images taken in that location.
    So big cities, which are the subjects of many images, have a more fine-grained grid structure than more remote regions where photographs are less common. Indeed, the Google team ignored areas like oceans and the polar regions, where few photographs have been taken.

  • Next, the team created a database of geolocated images from the Web and used the location data to determine the grid square in which each image was taken. This data set is huge, consisting of 126 million images along with their accompanying Exif location data.
  • Weyand and co used 91 million of these images to teach a powerful neural network to work out the grid location using only the image itself. Their idea is to input an image into this neural net and get as the output a particular grid location or a set of likely candidates. 
  • They then validated the neural network using the remaining 34 million images in the data set. 
  • Finally they tested the network—which they call PlaNet—in a number of different ways to see how well it works.
The results make for interesting reading. To measure the accuracy of their machine, they fed it 2.3 million geotagged images from Flickr to see whether it could correctly determine their location. “PlaNet is able to localize 3.6 percent of the images at street-level accuracy and 10.1 percent at city-level accuracy,” say Weyand and co. What’s more, the machine determines the country of origin in a further 28.4 percent of the photos and the continent in 48.0 percent of them.

That’s pretty good. But to show just how good, Weyand and co put PlaNet through its paces in a test against 10 well-traveled humans. For the test, they used an online game that presents a player with a random view taken from Google Street View and asks him or her to pinpoint its location on a map of the world.

Anyone can play at www.geoguessr.com. Give it a try—it’s a lot of fun and more tricky than it sounds.
GeoGuesser Screen Capture Example

Needless to say, PlaNet trounced the humans. “In total, PlaNet won 28 of the 50 rounds with a median localization error of 1131.7 km, while the median human localization error was 2320.75 km,” say Weyand and co. “[This] small-scale experiment shows that PlaNet reaches superhuman performance at the task of geolocating Street View scenes.

An interesting question is how PlaNet performs so well without being able to use the cues that humans rely on, such as vegetation, architectural style, and so on. But Weyand and co say they know why: "We think PlaNet has an advantage over humans because it has seen many more places than any human can ever visit and has learned subtle cues of different scenes that are even hard for a well-traveled human to distinguish.

They go further and use the machine to locate images that do not have location cues, such as those taken indoors or of specific items. This is possible when images are part of albums that have all been taken at the same place. The machine simply looks through other images in the album to work out where they were taken and assumes the more specific image was taken in the same place.

That’s impressive work that shows deep neural nets flexing their muscles once again. Perhaps more impressive still is that the model uses a relatively small amount of memory unlike other approaches that use gigabytes of the stuff. “Our model uses only 377 MB, which even fits into the memory of a smartphone,” say Weyand and co.

That’s a tantalizing idea—the power of a superhuman neural network on a smartphone. It surely won’t be long now!

Ref: arxiv.org/abs/1602.05314 : PlaNet—Photo Geolocation with Convolutional Neural Networks

ORIGINAL: TechnoplogyReview
by Emerging Technology from the arXiv
February 24, 2016

martes, 21 de febrero de 2017

End-to-End Deep Learning for Self-Driving Cars

In a new automotive application, we have used convolutional neural networks (CNNs) to map the raw pixels from a front-facing camera to the steering commands for a self-driving car. This powerful end-to-end approach means that with minimum training data from humans, the system learns to steer, with or without lane markings, on both local roads and highways. The system can also operate in areas with unclear visual guidance such as parking lots or unpaved roads.

Figure 1: NVIDIA’s self-driving car in action.
We designed the end-to-end learning system using an NVIDIA DevBox running Torch 7 for training. An NVIDIA DRIVETM PX self-driving car computer, also with Torch 7, was used to determine where to drive—while operating at 30 frames per second (FPS). The system is trained to automatically learn the internal representations of necessary processing steps, such as detecting useful road features, with only the human steering angle as the training signal. We never explicitly trained it to detect, for example, the outline of roads. In contrast to methods using explicit decomposition of the problem, such as lane marking detection, path planning, and control, our end-to-end system optimizes all processing steps simultaneously.

We believe that end-to-end learning leads to better performance and smaller systems. Better performance results because the internal components self-optimize to maximize overall system performance, instead of optimizing human-selected intermediate criteria, e. g., lane detection. Such criteria understandably are selected for ease of human interpretation which doesn’t automatically guarantee maximum system performance. Smaller networks are possible because the system learns to solve the problem with the minimal number of processing steps.

This blog post is based on the NVIDIA paper End to End Learning for Self-Driving Cars. Please see the original paper for full details.

Convolutional Neural Networks to Process Visual Data
CNNs[1] have revolutionized the computational pattern recognition process[2]. Prior to the widespread adoption of CNNs, most pattern recognition tasks were performed using an initial stage of hand-crafted feature extraction followed by a classifier. The important breakthrough of CNNs is that features are now learned automatically from training examples. The CNN approach is especially powerful when applied to image recognition tasks because the convolution operation captures the 2D nature of images. By using the convolution kernels to scan an entire image, relatively few parameters need to be learned compared to the total number of operations.

While CNNs with learned features have been used commercially for over twenty years [3], their adoption has exploded in recent years because of two important developments.

  • First, large, labeled data sets such as the ImageNet Large Scale Visual Recognition Challenge (ILSVRC)[4] are now widely available for training and validation. 
  • Second, CNN learning algorithms are now implemented on massively parallel graphics processing units (GPUs), tremendously accelerating learning and inference ability.
The CNNs that we describe here go beyond basic pattern recognition. We developed a system that learns the entire processing pipeline needed to steer an automobile. The groundwork for this project was actually done over 10 years ago in a Defense Advanced Research Projects Agency (DARPA) seedling project known as DARPA Autonomous Vehicle (DAVE)[5], in which a sub-scale radio control (RC) car drove through a junk-filled alley way. DAVE was trained on hours of human driving in similar, but not identical, environments. The training data included video from two cameras and the steering commands sent by a human operator.

In many ways, DAVE was inspired by the pioneering work of Pomerleau[6], who in 1989 built the Autonomous Land Vehicle in a Neural Network (ALVINN) system. ALVINN is a precursor to DAVE, and it provided the initial proof of concept that an end-to-end trained neural network might one day be capable of steering a car on public roads. DAVE demonstrated the potential of end-to-end learning, and indeed was used to justify starting the DARPA Learning Applied to Ground Robots (LAGR) program[7], but DAVE’s performance was not sufficiently reliable to provide a full alternative to the more modular approaches to off-road driving. (DAVE’s mean distance between crashes was about 20 meters in complex environments.)

About a year ago we started a new effort to improve on the original DAVE, and create a robust system for driving on public roads. The primary motivation for this work is to avoid the need to recognize specific human-designated features, such as lane markings, guard rails, or other cars, and to avoid having to create a collection of “if, then, else” rules, based on observation of these features. We are excited to share the preliminary results of this new effort, which is aptly named: DAVE–2.



The DAVE-2 System

Figure 2: High-level view of the data collection system.
Figure 2 shows a simplified block diagram of the collection system for training data of DAVE-2. Three cameras are mounted behind the windshield of the data-acquisition car, and timestamped video from the cameras is captured simultaneously with the steering angle applied by the human driver. The steering command is obtained by tapping into the vehicle’s Controller Area Network (CAN) bus. In order to make our system independent of the car geometry, we represent the steering command as 1/r, where r is the turning radius in meters. We use 1/r instead of r to prevent a singularity when driving straight (the turning radius for driving straight is infinity). 1/r smoothly transitions through zero from left turns (negative values) to right turns (positive values).

Training data contains single images sampled from the video, paired with the corresponding steering command (1/r). Training with data from only the human driver is not sufficient; the network must also learn how to recover from any mistakes, or the car will slowly drift off the road. The training data is therefore augmented with additional images that show the car in different shifts from the center of the lane and rotations from the direction of the road.

The images for two specific off-center shifts can be obtained from the left and the right cameras. Additional shifts between the cameras and all rotations are simulated through viewpoint transformation of the image from the nearest camera. Precise viewpoint transformation requires 3D scene knowledge which we don’t have, so we approximate the transformation by assuming all points below the horizon are on flat ground, and all points above the horizon are infinitely far away. This works fine for flat terrain, but for a more complete rendering it introduces distortions for objects that stick above the ground, such as cars, poles, trees, and buildings. Fortunately these distortions don’t pose a significant problem for network training. The steering label for the transformed images is quickly adjusted to one that correctly steers the vehicle back to the desired location and orientation in two seconds.

Figure 3: Training the neural network.
Figure 3 shows a block diagram of our training system. Images are fed into a CNN that then computes a proposed steering command. The proposed command is compared to the desired command for that image, and the weights of the CNN are adjusted to bring the CNN output closer to the desired output. The weight adjustment is accomplished using back propagation as implemented in the Torch 7 machine learning package.

Once trained, the network is able to generate steering commands from the video images of a single center camera. Figure 4 shows this configuration.
Figure 4: The trained network is used to generate steering commands from a single front-facing center camera.
Data Collection
Training data was collected by driving on a wide variety of roads and in a diverse set of lighting and weather conditions. We gathered surface street data in central New Jersey and highway data from Illinois, Michigan, Pennsylvania, and New York. Other road types include two-lane roads (with and without lane markings), residential roads with parked cars, tunnels, and unpaved roads. Data was collected in clear, cloudy, foggy, snowy, and rainy weather, both day and night. In some instances, the sun was low in the sky, resulting in glare reflecting from the road surface and scattering from the windshield.

The data was acquired using either our drive-by-wire test vehicle, which is a 2016 Lincoln MKZ, or using a 2013 Ford Focus with cameras placed in similar positions to those in the Lincoln. Our system has no dependencies on any particular vehicle make or model. Drivers were encouraged to maintain full attentiveness, but otherwise drive as they usually do. As of March 28, 2016, about 72 hours of driving data was collected.
Network Architecture

Figure 5: CNN architecture. The network has about 27 million connections and 250 thousand parameters.
We train the weights of our network to minimize the mean-squared error between the steering command output by the network, and either the command of the human driver or the adjusted steering command for off-center and rotated images (see “Augmentation”, later). Figure 5 shows the network architecture, which consists of 9 layers, including a normalization layer, 5 convolutional layers, and 3 fully connected layers. The input image is split into YUV planes and passed to the network.

The first layer of the network performs image normalization. The normalizer is hard-coded and is not adjusted in the learning process. Performing normalization in the network allows the normalization scheme to be altered with the network architecture, and to be accelerated via GPU processing.

The convolutional layers are designed to perform feature extraction, and are chosen empirically through a series of experiments that vary layer configurations. We then use strided convolutions in the first three convolutional layers with a 2×2 stride and a 5×5 kernel, and a non-strided convolution with a 3×3 kernel size in the final two convolutional layers.

We follow the five convolutional layers with three fully connected layers, leading to a final output control value which is the inverse-turning-radius. The fully connected layers are designed to function as a controller for steering, but we noted that by training the system end-to-end, it is not possible to make a clean break between which parts of the network function primarily as feature extractor, and which serve as controller.

Training Details

DATA SELECTION
The first step to training a neural network is selecting the frames to use. Our collected data is labeled with road type, weather condition, and the driver’s activity (staying in a lane, switching lanes, turning, and so forth). To train a CNN to do lane following, we simply select data where the driver is staying in a lane, and discard the rest. We then sample that video at 10 FPS because a higher sampling rate would include images that are highly similar, and thus not provide much additional useful information. To remove a bias towards driving straight the training data includes a higher proportion of frames that represent road curves.

AUGMENTATION
After selecting the final set of frames, we augment the data by adding artificial shifts and rotations to teach the network how to recover from a poor position or orientation. The magnitude of these perturbations is chosen randomly from a normal distribution. The distribution has zero mean, and the standard deviation is twice the standard deviation that we measured with human drivers. Artificially augmenting the data does add undesirable artifacts as the magnitude increases (as mentioned previously).

Simulation
Before road-testing a trained CNN, we first evaluate the network’s performance in simulation. Figure 6 shows a simplified block diagram of the simulation system, and Figure 7 shows a screenshot of the simulator in interactive mode.
Figure 6: Block-diagram of the drive simulator.
The simulator takes prerecorded videos from a forward-facing on-board camera connected to a human-driven data-collection vehicle, and generates images that approximate what would appear if the CNN were instead steering the vehicle. These test videos are time-synchronized with the recorded steering commands generated by the human driver.

Since human drivers don’t drive in the center of the lane all the time, we must manually calibrate the lane’s center as it is associated with each frame in the video used by the simulator. We call this position the “ground truth”.

The simulator transforms the original images to account for departures from the ground truth. Note that this transformation also includes any discrepancy between the human driven path and the ground truth. The transformation is accomplished by the same methods as described previously.

The simulator accesses the recorded test video along with the synchronized steering commands that occurred when the video was captured. The simulator sends the first frame of the chosen test video, adjusted for any departures from the ground truth, to the input of the trained CNN, which then returns a steering command for that frame. The CNN steering commands as well as the recorded human-driver commands are fed into the dynamic model [7] of the vehicle to update the position and orientation of the simulated vehicle.
Figure 7: Screenshot of the simulator in interactive mode. See text for explanation of the performance metrics. The green area on the left is unknown because of the viewpoint transformation. The highlighted wide rectangle below the horizon is the area which is sent to the CNN.
The simulator then modifies the next frame in the test video so that the image appears as if the vehicle were at the position that resulted by following steering commands from the CNN. This new image is then fed to the CNN and the process repeats.

The simulator records the off-center distance (distance from the car to the lane center), the yaw, and the distance traveled by the virtual car. When the off-center distance exceeds one meter, a virtual human intervention is triggered, and the virtual vehicle position and orientation is reset to match the ground truth of the corresponding frame of the original test video.

Evaluation
We evaluate our networks in two steps: first in simulation, and then in on-road tests.

In simulation we have the networks provide steering commands in our simulator to an ensemble of prerecorded test routes that correspond to about a total of three hours and 100 miles of driving in Monmouth County, NJ. The test data was taken in diverse lighting and weather conditions and includes highways, local roads, and residential streets.

We estimate what percentage of the time the network could drive the car (autonomy) by counting the simulated human interventions that occur when the simulated vehicle departs from the center line by more than one meter. We assume that in real life an actual intervention would require a total of six seconds: this is the time required for a human to retake control of the vehicle, re-center it, and then restart the self-steering mode. We calculate the percentage autonomy by counting the number of interventions, multiplying by 6 seconds, dividing by the elapsed time of the simulated test, and then subtracting the result from 1:


Thus, if we had 10 interventions in 600 seconds, we would have an autonomy value of


ON-ROAD TESTS
After a trained network has demonstrated good performance in the simulator, the network is loaded on the DRIVE PX in our test car and taken out for a road test. For these tests we measure performance as the fraction of time during which the car performs autonomous steering. This time excludes lane changes and turns from one road to another. For a typical drive in Monmouth County NJ from our office in Holmdel to Atlantic Highlands, we are autonomous approximately 98% of the time. We also drove 10 miles on the Garden State Parkway (a multi-lane divided highway with on and off ramps) with zero intercepts.

Here is a video of our test car driving in diverse conditions.


Visualization of Internal CNN State
Figure 8: How the CNN “sees” an unpaved road. Top: subset of the camera image sent to the CNN. Bottom left: Activation of the first layer feature maps. Bottom right: Activation of the second layer feature maps. This demonstrates that the CNN learned to detect useful road features on its own, i. e., with only the human steering angle as training signal. We never explicitly trained it to detect the outlines of roads.
Figures 8 and 9 show the activations of the first two feature map layers for two different example inputs, an unpaved road and a forest. In case of the unpaved road, the feature map activations clearly show the outline of the road while in case of the forest the feature maps contain mostly noise, i. e., the CNN finds no useful information in this image.

This demonstrates that the CNN learned to detect useful road features on its own, i. e., with only the human steering angle as training signal. We never explicitly trained it to detect the outlines of roads, for example.
Figure 9: Example image with no road. The activations of the first two feature maps appear to contain mostly noise, i. e., the CNN doesn’t recognize any useful features in this image.
Conclusions
We have empirically demonstrated that CNNs are able to learn the entire task of lane and road following without manual decomposition into road

  • Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1(4):541–551, Winter 1989.
    URL: http://yann.lecun.org/exdb/publis/pdf/lecun-89e.pdf
  • Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks
  • In F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 25, pages 1097–1105. Curran Associates, Inc., 2012. URL: http://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf.
  • L. D. Jackel, D. Sharman, Stenard C. E., Strom B. I., , and D Zuckert. Optical character recognition for self-service banking. AT&T Technical Journal, 74(1):16–24, 1995.
  • Large scale visual recognition challenge (ILSVRC). URL: http://www.image-net.org/challenges/LSVRC/.
  • Net-Scale Technologies, Inc. Autonomous off-road vehicle control using end-to-end learning, July 2004. Final technical report. URL: http://net-scale.com/doc/net-scale-dave-report.pdf.
  • Dean A. Pomerleau. ALVINN, an autonomous land vehicle in a neural network. Technical report, Carnegie Mellon University, 1989.
    URL: http://repository.cmu.edu/cgi/viewcontent.cgi?article=2874&context=compsci.
  • Danwei Wang and Feng Qi. Trajectory planning for a four-wheel-steering vehicle. In Proceedings of the 2001 IEEE International Conference on Robotics & Automation, May 21–26 2001. URL: http://www.ntu.edu.sg/home/edwwang/confpapers/wdwicar01.pdf.
    rlane marking detection, semantic abstraction, path planning, and control. A small amount of training data from less than a hundred hours of driving was sufficient to train the car to operate in diverse conditions, on highways, local and residential roads in sunny, cloudy, and rainy conditions. 
  • The CNN is able to learn meaningful road features from a very sparse training signal (steering alone).
  • The system learns for example to detect the outline of a road without the need of explicit labels during training.
  • More work is needed to improve the robustness of the network, to find methods to verify the robustness, and to improve visualization of the network-internal processing steps.
For full details please see the paper that this blog post is based on, and please contact us if you would like to learn more about NVIDIA’s autonomous vehicle platform!

REFERENCES

  1. Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backprop- agation applied to handwritten zip code recognition. Neural Computation, 1(4):541–551, Winter 1989. URL: http://yann.lecun.org/exdb/publis/pdf/lecun-89e.pdf.
  2. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, editors, Advances in Neural Information Processing Systems 25, pages 1097–1105. Curran Associates, Inc., 2012. URL: http://papers.nips.cc/paper/ 4824-imagenet-classification-with-deep-convolutional-neural-networks. pdf.
  3. L. D. Jackel, D. Sharman, Stenard C. E., Strom B. I., , and D Zuckert. Optical character recognition for self-service banking. AT&T Technical Journal, 74(1):16–24, 1995.
  4. Large scale visual recognition challenge (ILSVRC). URL: http://www.image-net.org/ challenges/LSVRC/.
  5. Net-Scale Technologies, Inc. Autonomous off-road vehicle control using end-to-end learning, July 2004. Final technical report. URL: http://net-scale.com/doc/net-scale-dave-report.pdf.
  6. Dean A. Pomerleau. ALVINN, an autonomous land vehicle in a neural network. Technical report, Carnegie Mellon University, 1989. URL: http://repository.cmu.edu/cgi/viewcontent. cgi?article=2874&context=compsci.
  7. Danwei Wang and Feng Qi. Trajectory planning for a four-wheel-steering vehicle. In Proceedings of the 2001 IEEE International Conference on Robotics & Automation, May 21–26 2001. URL: http: //www.ntu.edu.sg/home/edwwang/confpapers/wdwicar01.pdf.

ORIGINAL: NVidia