What is computer vision? | IBM

Applications of AI


Scientists and engineers have been trying to develop mechanical methods for viewing and understanding visual data for nearly 60 years. The experiment began in 1959 when neurophysiologists showed a series of images on cats, attempting to correlate brain responses. They found it first reacted to a hard edge or line and reacted scientifically. This means that image processing starts with simple shapes like straight edges.2

At about the same time, the first computer image scanning technology was developed, allowing computers to digitize and capture images. Another milestone reached in 1963, when computers were able to convert two-dimensional images into three-dimensional formats. In the 1960s, AI emerged as an academic learning field, and also marked the beginning of AI quests to solve human visual problems.

In 1974, Optical Character Recognition (OCR) technology was introduced, which allowed the recognition of text printed on fonts or typefaces.3 Similarly, intelligent character recognition (ICR) can decipher handwritten text using neural networks.4 Since then, OCR and ICR have found their way to document and invoice processing, vehicle plate recognition, mobile payments, machine conversion, and other popular applications.

In 1982, neuroscientist David Marr established that vision works hierarchically and introduces mechanical algorithms to detect edges, corners, curves and similar basic shapes. At the same time, computer scientist Fukushima Kiyoshi developed a network of cells that can recognize patterns. A network called NeoCognitron contains a convolutional layer in a neural network.

By 2000, the focus of research was on object recognition. And by 2001, the first real-time facial recognition application had arrived. Standardization of methods in which visual datasets are tagged and annotated emerged throughout the 2000s. In 2010, Imagenet datasets became available. It contains millions of tagged images in a thousand object classes, providing the foundation for the currently used CNNS and deep learning models. In 2012, a University of Toronto team took part in an image recognition contest at CNN. A model called AlexNet has significantly reduced image recognition error rates. After this breakthrough, the error rate fell to just a few percent.5



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *