Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Lecture 57 min.



The problem, errors and difficulties of pattern recognition are among the central issues in the concept of artificial intelligence. This article does not discuss the image recognition algorithms themselves, but rather the limits of their applicability, in particular when creating artificial intelligence.

In my opinion, visual pattern recognition by a human and by a computer system differ so greatly that they have little in common. When a person says "I see," they are in fact thinking more than seeing, which cannot be said of a computer system equipped with image recognition hardware.

I know the idea is not new, but let us once again verify that it is correct using the example of a robot that claims to possess intelligence. The test question is this: how must a robot see the world around it in order to become completely like a human?

Of course, the robot must recognize objects. Oh yes, the algorithms cope with that, by learning from initial samples, as I understand it. But that is catastrophically little!

Classical approaches to pattern recognition

In this article, "classical" refers to approaches or methods of pattern recognition that satisfy two conditions:

  • these methods must be used within the "classical" pattern recognition scheme;
  • these are the most widespread pattern recognition methods over the last decades.

"Classical" refers to the pattern recognition problem in the following formulation (see the foreword by Yu.I. Zhuravlev to the work ): "the objects to be recognized are specified by a set of features, a certain number of reference objects are known, and their descriptions make up the initial (training) information. On the basis of this information, an algorithm is synthesized that determines, for newly arriving objects, which of a finite number of classes (or which ones) they belong to." The same recognition scheme is given in the reference book on artificial intelligence [5, p. 150].

It seems that the main limitations of classical pattern recognition methods can be expressed by the following statements:

  • the incompleteness of the classical pattern recognition scheme, and the absence in it of mechanisms for its development and its embedding into the scheme of cognition;
  • the insufficiency and inconsistency of existing approaches to the content of the concept of information, and the absence in them of any possibility of representing the transfer of information between levels during recognition;
  • the incompleteness, imprecision and inconsistency of the concept of "pattern" and its components (feature, etc.);
  • the inadequacy of the classical scheme when using the concepts of "pattern" and "class";
  • insufficient use of a priori information. The insufficiency of discrete mathematics and of the generally accepted computational paradigm;
  • insufficient account of the multilevel nature of information processing;
  • the impossibility of creating a full-fledged theory of recognition within a single scientific direction, even the most promising one.

1. Incompleteness of the classical pattern recognition scheme

The classical recognition scheme, which uses the concepts of features, their combination and subsequent classification, relies on several implicit, never declared assumptions, which boil down to the following:

at any given moment only one pattern is perceived, and either the level of information related to the background is insignificant, or we have simple tools for separating the part of the information related to the background from the information about the features;

the task of the perceiving subject (a human or a robot) is reduced to assigning an object to one of the classes and nothing more.

In reality, neither of these assumptions holds. We (or a robot) perceive the surrounding world, which contains many objects. Depending on the tasks, we single out some of these objects and perceive others as background, and the recognition task in this case is considerably more complex than the classical scheme: a) we must extract from the noise the information about the features of the set of objects to be recognized, and it is not obvious in advance what in the signal should be considered noise and what useful information, i.e. which features to extract from the noise. Obviously, this task has many possible solutions and we must choose the right one; b) we must somehow assign the features to some object, and even without numerical estimates it is obvious that this task cannot be solved combinatorially, and here it is necessary to use certain principles that sharply limit the number of possible variants; c) we must then recognize these objects, single out those of interest to us and relate them to our problems.

Regarding the second assumption: pattern recognition is not an independent procedure, it is built into the scheme of perception and thinking. And in this scheme recognition makes up only part of the task. The second part of the task is to pick out, against the background of some relatively invariant part, the features of the pattern or situation that characterize it here and now, and to use these specific features in thinking. For example, when we meet an old acquaintance, we do not simply assign him to one of the classes, but determine his mood, learn about his problems, and so on. And at the end of the meeting we do not assign the new pattern we have obtained to some existing class – this is impossible for purely combinatorial reasons – but construct a new pattern consisting of two parts: some invariant I and a unique variable part delta, which we compose from elements of our impressions and language (essentially, concepts, which can also be represented as a certain class of patterns). Thus, the result of perception V with respect to an individual pattern singled out by us is more correctly expressed by the formula:

V = I + delta.

Let us stress the fundamental importance of representing patterns in this form. The work of thinking and intelligence is largely determined precisely by the unique, not the invariant, part of perceived patterns, so by not taking it into account we essentially throw away the most important

Let us stress the fundamental importance of representing patterns in this form. The work of thinking and intelligence is largely determined precisely by the unique, not the invariant, part of perceived patterns, so by not taking it into account we essentially throw away the most important components of thinking and intelligence. It becomes obvious that for this reason alone (the neglect of the unique component of the pattern) the classical recognition scheme cannot be used to model the work of intelligence. Neglecting the unique component of the pattern leads to the so-called algorithmic or automaton approach, which is equivalent to the work of the nervous system at the level of an insect. It seems that most specialists working on pattern recognition would agree that the artificial recognition systems developed to date reach, at best, the level of an insect. Note that neither cybernetics as a science nor pattern recognition has managed to overcome the automaton barrier, as a result of which information technologies, Data Mining and a number of other new disciplines appeared in place of cybernetics, while the science of pattern recognition is experiencing a crisis. (Unfortunately, these new disciplines have not been able to fully compensate for the decline of cybernetics, since they do not properly use the systems principle underlying cybernetics – and this is a weighty argument in favor of a renewed cybernetics).

2. Insufficiency and inconsistency of the concept of information

It is well known that the recognition process in humans and animals is a complex, multistage process, accompanied by repeated recoding, compression, generalization and change in the representation (the language and the set of features) of the information related to the pattern, or, more precisely, to the patterns together with the background and noise (experimental data on this question are set out quite well in , and their analysis is carried out, in particular, in [10]). It seems natural that the description of all these procedures should be carried out on the basis of the concept of information. But does the modern concept of information allow this? One can state unambiguously and categorically – no, it does not. Let us briefly consider the reasons. Only the laziest has not scolded the concept of information; the literature contains many differing formulations of this concept, but we will note only the best known ones, which are customarily used in serious research. In 1948, C. Shannon defined the amount of transmitted information as the logarithm of the number of possible states of a signal [11] (the combinatorial approach). The introduced concept became widely known and was modified in several directions. First, a close connection was shown between information understood in this way and entropy. Second, the combinatorial approach was supplemented by a probabilistic interpretation (while the formulas for the amount of information retained practically the same form). And third, the concept of conditional probability was introduced as the difference of the logarithms of the states or choices of the system before and after receiving the information. Later, A.N. Kolmogorov [12] proposed an algorithmic interpretation of the amount of information as the shortest description of a program for transforming an object from one state into another. All these approaches to defining the concept of information share two common features:

1) it is assumed that both the message as a whole and all its component parts have informational value (and the same value) for the perceiving system, i.e. each part of the signal reduces the number of possible states of the system;

2) it is assumed that all states or choices of the perceiving system are equally probable.

In reality neither of these assumptions holds, and therefore in practice we mostly use a different understanding of information, based not on the length of the message but on its importance to the perceiving subject. In this connection a large number of approaches to information based on its importance have been proposed (see, e.g., the review of the state of this question in [13]). But none of these approaches has gained wide recognition at present. This is apparently because the concepts of value, importance and the like cannot be introduced without subjectivism – dependence on the subject, the task, the complexity of the information, the stage of its processing, etc. In this case we obtain a multicriteria problem. But it is well known that problems of this type generally have many possible solutions (the Pareto set). In that case, we are hardly likely in the foreseeable future to have a value-based interpretation of information that is more clearly defined than, for example, the concepts of freedom, justice or morality. Thus, none of the existing interpretations of information allows this concept to be used to describe the process of transferring information between levels and representations (languages) in the course of pattern recognition. In this connection it may be necessary to develop concepts that replace the concept of information but are narrower in content and allow a more unambiguous and adequate description of information transformations in the course of recognition. In addition, it is necessary to use and develop the theory of translation. The emphasis here should be placed on understanding the essence of translation between languages of different (rather than the same) level and different representations.

3. Inadequacy of the classical scheme in its use of the concept of pattern and its components

In the process of recognition according to the classical scheme, several assumptions (they can also be called principles, approaches, schemes, etc.) are made regarding the pattern being recognized, in particular the following: – on the way between the initial "raw" information and the final recognized pattern, some intermediate stage of ordering the initial information is introduced, called a feature or a set of features; – recognition is carried out according to the scheme: obtaining an image, extracting features from it, obtaining a pattern as a combination of features, the process of assigning the pattern to one of the classes (let us call this bottom-up recognition); – in all cases both the set of recognized features and the pattern obtained from it can be recognized in only one correct way. How well founded are these assumptions? Let us consider the first assumption. As follows from the literature, "the task of feature selection… is understood as a transformation (mapping) of the initial space into a feature space of lower dimension without loss of information" [14, p. 8], i.e. with the help of features, at the first stage, both the description of the pattern and the subsequent operation on information about the pattern are substantially reduced. At the second stage the features are used to build the pattern in the form of a set of features, and then recognition proper takes place. While the second stage is lovingly, carefully and in detail described by numerous authors, little is written about the first, and mostly in terms of goals rather than algorithms. The unsatisfactory state of this question is noted, in particular, in [2, p. 152]: "We have learned well how to solve problems in which the feature space has already been built one way or another before training, or the vocabulary of the language is given… And the more successfully the features are chosen (by a human, not by a machine), the more compact the patterns are, and hence the more successfully the learning problem is solved… But if the features are not specified in advance or are chosen randomly and unsuccessfully, all methods lose their working capacity, and their reliability and quality drop catastrophically… So, we have not yet learned the main thing – automating the construction of such spaces of images of objects in which the objects would be compact."

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Example

The autopilot's recognition system mistook the Moon for a traffic light showing a yellow signal and began to slow the car down.

To understand the reasons for the insufficient attention to the problem of feature extraction, let us consider pattern recognition at the conceptual level in humans. Obviously, the recognition process in this case consists, first, of more than two levels, and second, at all the higher levels we work, in essence, not with features but with concepts, which are more correctly represented as patterns rather than as features. For example, when we recognize a person, according to the classical scheme we first recognize features. In particular, let one of these features be glasses. But the pattern of glasses is quite complex in itself, and in another recognition task glasses may act as an independent pattern to be recognized, containing subpatterns, for example, lenses, etc. Consequently, at least at the level of conceptual thinking, the recognition scheme looks not like a feature – pattern chain, but like a level-1 pattern – level-2 pattern … – final pattern chain. Thus, introducing the concept of a feature is a kind of intellectual trick (at least at the level of conceptual thinking), with the help of which we throw recognition proper out of consideration of the recognition process and leave the more tractable task of manipulating patterns of various levels (it is more tractable when we throw out of this scheme the outputs into thinking, i.e. the variable part of the pattern, see section 1). These arguments are apparently less applicable to simpler cases of recognition, for example, of speech or of simple images. This is because in this case the features, or an intermediate stage on the way to the features, are fairly simple invariants (primitives, in the terms of D. Marr [15]), for which algorithms have been developed for obtaining them from the initial "raw" information.

Now let us consider the recognition sequence: initial information – feature – pattern. As indicated above, following the logic of the existing approach to recognition, for complex cases it is more correctly expressed by the chain: initial information – level-1 pattern… – final pattern. But if we proceed from experimental data, then in humans, at least in visual perception, a completely different scheme works. In this scheme our recognition of a scene as a whole is based first of all on a certain holistic understanding, which in psychology is called a gestalt. "The main empirical essence of the phenomenon of the wholeness of the perceptual image discovered by Gestalt psychology… lies in the dominance of the holistic structure of the percept over the perception of its individual elements" [16, ch. 9]. An objection may follow that these conclusions relate to the process of thinking – the stage following recognition, when we refine the content of the scene. However, there are weighty neurophysiological data indicating that the process of pattern recognition itself is carried out according to the gestalt scheme. Thus, it was shown in [17] that different spatial-frequency components of a complex image in living beings are extracted and described in the visual cortex by independently working channels. Later, in the works of V.D. Glezer [18] and other researchers, a scheme using the Fourier transform was proposed and experimentally substantiated. The low-frequency Fourier transform corresponds to the large-scale details of the image, and the high-frequency part to the fine details.

It has been shown experimentally that people perceive large-scale details faster than small-scale ones (for a detailed review and analysis of this question see [18, ch. 3]). Accordingly, recognition begins with the large details that describe the scene as a whole, rather than with its individual features. This conclusion is a simple consequence of the laws of information transmission: low resolution requires less information than high resolution, and because of the low information transmission rate in humans (about 10^2 impulses/sec per nerve fiber) or the large volume of information, information about image details is received by the recognition system (the brain) noticeably later than the outline of the scene as a whole. Thus, human pattern recognition would be more correctly described by the following scheme: initial information – the image as a whole – refinement of the details (features) of the image (top-down recognition). In more complex cases, given that a person is able to operate with 7 ± 2 objects simultaneously, the recognition scheme apparently becomes more complicated and looks approximately as follows: initial information – simultaneous extraction and preliminary recognition of several images – simultaneous refinement of the details of each of the images and recognition of the complex image, or assigning meaning to the scene (a combination of top-down and bottom-up recognition elements). Taking into account the state of artificial and natural recognition systems, it appears that there is not and cannot be a single correct scheme of pattern recognition. A wide variety of schemes may be used, and this is determined by the subject area, the goals of the research and the available tools. The basis for choosing such schemes should be a general theory of transformation and transmission of information between representations and languages of different levels and origins, which is currently lacking. In conclusion of this topic, let us add that experimental data [17-19] show that sufficiently complex recognition systems, in particular in humans and higher animals, use several recognition schemes simultaneously. Let us now consider the assumption that the pattern or patterns to be recognized are unique. Obviously, this assumption relies substantially on the assumption that the classification is unique. Without disputing this assumption for some simple cases, we note that in complex cases it, as a rule, does not hold [20, ch. 2]. For example, specialists know well that in describing socio-economic processes, a wide variety of classification systems are possible depending on the goals set. Apparently, this conclusion can be extended to all complex systems and to all incomplete and inexact descriptions, that is, except for special cases, to practically everything that we use to describe and represent information at the conceptual level. In such cases the recognition task is supplemented by the stage of choosing a classification system. In fairness, it should be noted that even in such a perfect recognition system as the human being, the classification of complex information takes place not in the process of recognition but in the process of analyzing information, that is, thinking. Here we see once again that the boundary between recognition and thinking is very thin and that we cannot consider recognition without thinking. Moreover, the thinking process itself, it turns out, can also be built into the recognition process. Thus, in conclusion of this section, let us note that the assumptions of the classical recognition scheme given at the beginning of the section are either applicable only to the simplest special cases (use of a single classification), or work exactly the opposite way (the bottom-up recognition scheme), or lead to a substitution of the recognition problem (introducing the notion of a "feature" without procedures for obtaining it).

4. Insufficient Use of A Priori Information

To assess the importance of a priori information more clearly, let us imagine that we are trying to recognize an image without any preliminary knowledge of it. It is easy to see that the problem posed in this way is equivalent to the problem of establishing contact (understanding) with an extraterrestrial civilization on the basis of radio signals. Several decades ago such an idea (the possibility of understanding or recognition) was hotly discussed in society, but no unambiguous answer was ever obtained. This is not surprising, given the low level of, and the crisis in, the scientific field connected with recognition. But even in the foreseeable future, provided that all the problems listed above are eliminated, it is hard to suppose that the task of recognizing a message from an extraterrestrial civilization from a radio signal (but not in a process of dialogue) is feasible in full without taking into account some additional information. In fact, a priori information is, of course, used in the recognition process (otherwise the science of recognition would not exist in any form). In particular, the preliminary choice of features can be made only with the use of this additional, usually poorly formalized information about the subject area, and, consequently, features chosen in advance contain a priori information about the image [1, p. 23]. Some a priori information is also contained in the recognition schemes, the tools used, the form of the results obtained, and so on.

The second method of using a priori information is learning from examples. For solving a number of problems, learning is often sufficient. The problem, however, is that there is no general theory (method, way) of constructing optimal features on the basis of a priori information. Moreover, as already pointed out earlier, the real recognition process in living beings is multilevel in character, and in this case a theory of recognition does not exist at all. In a number of cases, additional information about the image to be recognized may be presented only in fragments, incompletely and inexactly. For all these cases there is no theory (procedures, mechanisms) that would make it possible, in a sufficiently simple way, to obtain the maximum effect from scattered additional information about the image to be recognized, presented in various forms. How can this be done? There are apparently no immediate and simple recipes here. The theory of information needs to be developed, and a deeper understanding is needed of the role and mechanisms of manifestation of context, external criteria and similar concepts [10], [21].

5. Inadequacy of Existing Mathematics and the Computation Paradigm

One of the key problems of pattern recognition is the need to carry out a large volume of computation, which often cannot be implemented even on supercomputers. For example, in statistical recognition methods the number of decision rules that must be estimated in the recognition process, even for the simplest linear binary rules, is of the order of 2^m, where m is the dimension of the space [1, p. 91], equal to the number of features. Given that the number of features in many problems exceeds tens and hundreds, we obtain problems that cannot be computed by traditional methods (the curse of dimensionality). There are several directions within which attempts are made to solve the computational problems. Conventionally they can be divided into three classes:

  • neural network methods;
  • new methods within the traditional ideology;
  • non-traditional methods (the brain and quantum computers).

The methods developed both within the traditional ideology and in neural networks explicitly or implicitly use the existence theorem proved by A.N. Kolmogorov on the representability of a function of many variables by a superposition of a finite number of functions of one variable (under a number of restrictions) [22]. From the point of view of the recognition problem, this theorem is important because it shows the theoretical possibility of a sharp reduction in the dimension of the space, as a result of which the volume of computation grows with the number of features not exponentially but polynomially. One of the widely known methods that makes substantial use of the principles laid down in Kolmogorov's theorem is, in our view, the Group Method of Data Handling (GMDH) proposed by A.G. Ivakhnenko ([23], [24], as well as other works on the site www.inf.kiev.ua/GMDH-home). Generally speaking, this method is not reduced to the indicated theorem alone, and its second significant positive feature is the presence of mechanisms for building a model of optimal complexity that take into account a short sample or noisy initial data. This method is now developing rapidly through a whole group of researchers from various countries, a commercial product based on it (KnowledgeMiner) has been created, and it is also important that it can be implemented both with traditional sequential algorithms and with neural networks. Thus, it has every prerequisite to become one of the leading approaches in the use of recognition systems and, in general, of systems for obtaining new knowledge. At the same time, it is necessary to note the insufficient analysis of the theoretical premises underlying it, which hampers its development and obscures the interpretation of the results obtained. Also very interesting appears the method of reduction, or the method of limiting simplifications, proposed by V.I. Vasiliev [2, ch. 8], [25], [26], based on the maximum possible reduction of the dimension of the feature space, as a result of which the requirements both for the sample size and for the amount of computation are sharply reduced.

In statistics this ideology (reducing the dimension of the space) is well known, thoroughly developed and widely used [27, ch. 12-17]. The limitation of this ideology is also known: we cannot discard important features. In the limiting case where all features are equally important, we cannot reduce the dimension of the feature space at all. Therefore it seems that the main area of applicability of reduction theory is not the recognition process itself in the classical scheme, but the stage of feature selection – the process of constructing the feature space. And since this process must be based on thinking, reduction theory is closely related to thinking, language, ways of representing information, and so on, and obviously cannot be developed in isolation from these concepts. Classical mathematics is based on the principle of sequential computation. In doing so, it implicitly assumes that we can represent the process of transforming information about any subject area as a sequential chain of computations (let us call it the principle of locality). But this is not the only principle of information processing possible at present. The principle of parallel processing, used in particular in neural networks, is also well known and widely used; it makes it possible to speed up information processing by a factor of n, where n is the number of parallel channels. But there is one more principle – the principle of joint information processing, when mutually consistent processing of information is carried out in a number of algorithms or elements of a computing device (nonlocal information processing). This principle is implemented, in particular, in quantum computers [28], [29] – a new, promising and rapidly developing direction. It is also possible that our brain works on the same principle, at least in some respects. The notion of "neural networks" unites hundreds of different methods with different ideologies and hardware implementations, which is why a survey of these methods from a unified standpoint, as already noted in the preface, is extremely difficult. For this reason we will only note that the most promising from the point of view of pattern recognition are considered to be Kohonen networks, the concept of the cognitron or neocognitron [30], [31] and the theory of invariants (see the collection of materials at www.kyb.tuebingen.mpg.de). One also cannot fail to mention attempts to overcome the Kolmogorov complexity barrier with the help of neural networks [32] (which also has a review of the literature on this question). If this attempt succeeds in full, the effect of it (at least in pattern recognition) will be comparable to the use of quantum computers. The operation of a quantum computer is based on controlled change of selected states in a quantum system consisting of n quantum particles – qubits. Such a quantum system has a number of remarkable features. First, the quantum states of the qubits influence each other, leading to the appearance of 2^n basis states. Second, if we switch one of the qubits to another state, then all the other states very quickly (at the speed of interference processes) go over to a new state, taking into account the change in the state of the original qubit. With some additional algorithms, such a process can be represented as a process of computation. And third, which is important, we can read out information about the new state, that is, obtain the results of the calculations [29]. If we consider this process within traditional concepts, it is not simply parallel computation (in parallel computation the individual processes are independent of each other, at least for some cycles), but mutually consistent information processing, in which each qubit constantly "feels" the states of the other qubits and constantly changes its own state consistently with them. As a result, the speed of calculation grows exponentially with the number of elements in a quantum computer. At present only individual elements of quantum computers have been implemented in practice, but specialists have no doubt that it is feasible in principle; disagreements mainly concern the timeframe – 5 or 10 years. The expected increase in the performance of quantum computers reaches, by some estimates, 10 or more orders of magnitude, which will undoubtedly make it possible to solve many recognition problems that are currently intractable. In connection with quantum computers one also cannot fail to mention the capabilities of the human brain. As is known, a person is able to hold 7 ± 2 objects in consciousness. Moreover, if we operate on one of the objects, more or less consistent changes take place in the others as well (connections, interpretation of features, the overall meaning of the given object for the situation as a whole, etc.) – all this is very reminiscent of the operation of a quantum computer. Consequently, one may suggest that the brain, at least with respect to some information-processing operations, carries out mutually consistent simultaneous processing of information. And this nonlocal processing is implemented not with the use of quantum computers, but through a complex structure of interneuronal connections and the transmission of impulses between neurons. Perhaps it is precisely the use of this principle that explains the enormous capabilities of the human brain, despite the fact that the number of elements in modern computers (at least in multiprocessor systems) is already closely approaching the number of neurons in the human brain.

6. Insufficient Consideration of the Multilevel Nature of Information Processing

Specialists in neurophysiology have shown experimentally that in humans and animals information processing is multilevel and multistage in character. Even at the preliminary stage of extracting information, when recognition is not yet under way, 3 – 4 independent levels are distinguished. In general, the number of levels, depending on the type of information processed, may well be 6 – 10. Each level has its own "hardware" part (regions of the brain), its own way of representing information and its own language. It has also been shown experimentally that, at least at the initial stage, information processing at each level is carried out by simple methods that are clear to us and can be implemented fairly easily at the hardware level or in software. At the same time, artificial recognition systems, as a rule, have only three levels (initial information – features – image) or, at best, four (initial information – non-derivative elements – features – image). Considering the superiority of living beings over artificial systems in pattern recognition, one can conclude that a multilevel representation of information is more effective.

7. The Impossibility of Creating a Complete Recognition Theory Within a Single Scientific Direction

Despite listing the limitations of the classical approach, the author by no means believes that there are no achievements in pattern recognition. On the contrary, over recent decades a number of powerful interdisciplinary directions have been developed, requiring significant intellectual and practical effort in neurophysiology, theoretical and applied mathematics, physics, artificial neural networks, and so on. These directions include the statistical theory of pattern recognition, the creation of a number of powerful computational methods, the discovery from experimental data of a number of important principles of information processing in humans, the development of the ideology of semantic networks, the theory and practice of neural networks... The list could be continued. An objection arises: something in this list turned out to be not quite adequate, but let us determine the most suitable method or ideology and, in the end, solve the recognition problem. It seems that this is the wrong approach.

As already noted in the preface, there are good reasons to believe that the key problems of pattern recognition, cybernetics and artificial intelligence share the same origin: we are trying to describe a subject that is too complex with a conceptual apparatus that is too simple. In particular, the subject whose operation has to be explained includes the knowing subject itself – the human brain. This task is difficult, considerably more difficult than quantum mechanics and, possibly, the most difficult that humanity has faced so far. Its difficulty lies, above all, in the need for a radical and, no less important, mutually consistent transformation of a whole range of scientific disciplines. Information theory must describe the transformation of information at all levels and stages of its processing – visual, linguistic, and so on; the laws of generalization, the principles of optimal representation of information, etc. must be established. Mathematics must have a fully developed theory of transforming complex multidimensional data presented in incomplete and inexact form, as well as through any kinds of description – formal, in natural languages, and so on. Neural networks must reach the level of modeling the most important principles of the operation of the human brain and of quantum computers.

8. Every object in the surrounding world consists of many other objects and is in turn a subset of still other objects.

This property can be called nesting. But what if some object simply has no name, and therefore is not in the database of source samples on which the algorithm is trained – what should the robot recognize in that case?

The cloud that I am observing through the window at this moment has no named parts, although it obviously consists of edges and a middle. However, special terms for the edges and the middle of a cloud do not exist; they have not been invented. To refer to an unnamed object I used a verbal formulation ("cloud" is the type of the object, "edge of the cloud" is the verbal formulation), which is beyond the capabilities of an image recognition algorithm.

It turns out that an algorithm without a logic block is good for very little. If the algorithm detects a part of a whole object, it will not always be able to figure out – and accordingly the robot will not be able to report – what it is.
Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects


9. The list of objects that make up the surrounding world is not closed: it is constantly being added to.



A human being is able to construct objects of reality by giving names to newly discovered objects, for example species of fauna. A horse with a human head and torso he will call a centaur, but first he will realize that this creature has a human head and torso while everything else is equine, and thereby recognize the observed object as a new one. This is how the human brain works. An algorithm, lacking input data, will identify such a creature either as a human or as a horse: since it does not operate with the characteristics of types, it will not be able to establish their combination.

For a robot to become like a human, it must be able to identify types of objects that are new to it and to give those types names. The descriptions of a new type must involve the characteristics of known types. And if the robot cannot do this, what on earth do we need such a fancy robot for?

Suppose we send a scout robot to Mars. The robot sees something unusual, but it can identify the object only in the terms of Earth that it knows. What will that give the people listening to the verbal messages coming from the robot? Sometimes it will give something, of course (if Earth-like objects turn up on Mars), and in other cases nothing (if Martian objects turn out not to resemble Earth ones).

An image is another matter: a person can see everything, assess it correctly and name it on their own. Only by means not of a pre-trained image recognition algorithm, but of their own, more cleverly built human brain.

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

10. There is a certain problem with the individualization of objects.



The surrounding world consists of concrete objects. Strictly speaking, only concrete objects can be seen. But in some cases they need to be individualized verbally, for which either personal names ("Vasya Petrov") or a simple pointing to a specific object, spoken or implied ("this table"), are used. What I call types of objects ("people", "tables") are merely collective names for objects that have certain common characteristics.

Image recognition algorithms, if trained on source samples, will be able to recognize both individualized and non-individualized objects – this is good. Face recognition in places where crowds gather and so on. The bad thing is that such algorithms will not understand which objects should be recognized as having individuality and which categorically should not.

A robot, as the owner of an AI, should from time to time burst out with statements such as:
– Oh, I saw this old lady a week ago!

But it would not do to overuse such remarks about blades of grass, especially since there are well-founded concerns about whether computing power is sufficient for such a task.

I do not understand where the fine line lies between the individualized old lady and the countless blades of grass in the field, which are individualized in themselves no less than the old lady, but which are of no interest to a human from the point of view of individualization. What is a recognized image in this sense? Almost nothing – the beginning of an agonizingly complex perception of the surrounding reality.

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Fig. This is one and the same person. Sergey Zverev

11. Fourth, the dynamics of objects, determined by their mutual spatial arrangement



I am sitting in a deep armchair in front of the fireplace and now I am trying to get up.
– What do you see, robot?

From our everyday point of view, the robot sees me getting up from the armchair. What should it answer? Probably the relevant answer would be:
– I see you getting up from the armchair.

For this the robot must know who I am, what an armchair is and what it means to get up…

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

An image recognition algorithm, after appropriate tuning, will be able to recognize me and the armchair, and then by comparing frames we can determine the fact of my moving away from the armchair, but what does "getting up" mean? How does "getting up" happen in physical reality at all?

If I have already got up and walked away, it is all fairly simple. After I moved away from the armchair, all the objects in the study have not changed their spatial positions relative to one another, except for me, who was originally in the armchair and after some time ended up at a distance from the armchair. It is permissible to conclude that I have left the armchair.

If I am still in the process of rising from the armchair, it is somewhat more complicated. I am still next to the armchair, but the mutual spatial position of the parts of my body has changed:

  • initially the shin and the torso were in a vertical position, and the thigh in a horizontal one (I was sitting),
  • at the next moment all the parts of the body were in a vertical position (I stood up).


If a person were observing my behavior, he would instantly conclude that I am getting up from the armchair. For a person this would be not so much a logical conclusion as a visual perception: he would literally see that I am getting up from the armchair, although in reality he would see a change in the mutual position of the parts of my body. However, in reality this will be a logical conclusion, which someone must explain to the robot, or which the robot must work out on its own.

Both are equally difficult:

  • entering into the initial knowledge base the information that getting up is a sequential change in the mutual spatial position of certain parts of the body is not particularly inspiring;
  • it is no less foolish to hope that the robot, as an artificial thinking being, will quickly guess on its own that the change described above in the mutual spatial position of certain parts of the body is called getting up. For a human this process takes years; how long will it take for a robot?


And what do image recognition algorithms have to do with it? They will never be able to determine that I am getting up from the armchair.


12. The "problem of nonstandard actions"

"Getting up" from the previous example is an abstract concept, defined by a change in the characteristics of material objects, in this case a change in their mutual spatial position. In general this holds for any abstract concepts, since abstract concepts by themselves do not exist in the material world but depend entirely on material objects. Although we often perceive them as if observed with our own eyes.

Moving the jaw to the right or to the left without opening the mouth – what is this action called? Nothing. Undoubtedly because such a movement is not really characteristic of humans. The robot will indeed see it with the algorithms under discussion, but what is the use? The required name will be absent from the database of source samples, and the robot will find it hard to name the recorded action. And image recognition algorithms are not trained to give extended verbal formulations to unnamed actions, or to other abstract concepts.

In essence, we have a duplicate of the first point, only concerning not objects but abstract concepts. However, the other points, both preceding and following, can also be linked with abstract concepts – I am simply drawing attention to the rise in the level of complexity when working with abstractions.


13. Taking cause-and-effect relations into account in pattern recognition



Imagine that you are watching a pickup truck fly off the road and knock down a fence. The cause of the fence being knocked down is the motion of the pickup, and in turn the motion of the pickup results in the fence being knocked down.

– I saw it with my own eyes!
This is the answer to the question of whether you saw what happened or figured it out. And what did you actually see?

Several objects, in this dynamic:

  • the pickup drove off the road,
  • the pickup came right up to the fence,
  • the fence changed shape and location.


Based on visual perception, the robot must realize that in the ordinary case fences do not change shape and location: here it happened as a result of contact with the pickup. The object that is the cause and the object that is the effect must be in contact with each other, otherwise there is no causality in their relationship.

Yet here we fall into a logical trap, because other objects, not only the causing object, can also be in contact with the object that is the effect.

Suppose that at the moment the pickup hit, a jackdaw landed on the fence. The pickup and the jackdaw were in contact with the fence simultaneously: how can we determine as a result of which contact the fence was knocked down?

Probably by means of repeatability:

  • if in every case when a jackdaw lands on a fence the fence is knocked down, the jackdaw is to blame;
  • if in every case when a pickup crashes into a fence – the pickup is to blame.

Recognition Errors for Nested, Transparent, Mirrored and Other Objects
Thus, the conclusion that the fence was knocked down by the pickup is not quite an observation, but the result of an analysis based on observing objects in contact with one another.

On the other hand, an influence can act at a distance, for example the action of a magnet on an iron object. How will the robot guess that the approach of a magnet makes a nail rush toward the magnet? The visual picture is not like that:

  • the magnet approaches but does not touch the nail,
  • at that very instant the nail, of its own accord, rushes toward the magnet and touches it.


As you can see, tracking cause-and-effect connections is very difficult, even in cases when a witness declares with iron conviction that he saw it with his own eyes. Image recognition algorithms are powerless here.



14. The choice of goals of visual perception and the problem of nested objects.



The surrounding visual picture can consist of hundreds and thousands of objects nested within one another, many of which are constantly changing their spatial position and other characteristics. Obviously the robot has no need to perceive every blade of grass in a field, nor, for that matter, every face on a city street: only what is important, depending on the tasks being performed, needs to be perceived.

Obviously, it will not work to tune an image recognition algorithm to perceive some objects and ignore others, since it may not be known in advance what to pay attention to and what to ignore, especially as current goals may change along the way. A situation may arise in which it is first necessary to perceive many thousands of objects nested within one another – literally every one of them – analyze them, and only then deliver a verdict on which objects are essential for the current task and which are of no interest. This is exactly how a human perceives the surrounding world: he sees only what is important, paying no attention to uninteresting background events. How he manages to do this remains a mystery.

And the robot, even one equipped with the most modern and ingenious image recognition algorithms?.. If during an attack by Martian aliens it starts its report with a weather summary and continues with a description of the new landscape spreading out before it, it may not manage to report the attack itself.

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects
15. The reconstruction problem in pattern recognition

Not all images can be entirely within the field of view, which creates problems for their recognition and identification

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

The photo shows the use of radio waves in the "RF-Pose" project to solve a problem of this kind; however, try to do this with an unpredictable object and without radio waves

15. The problem of recognizing translucent, transparent and mirror objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Objects made of transparent material or with a mirror coating are difficult to recognize and identify. This is a major problem for recognition algorithms because, as a rule, they determine the shape and boundary of an object by comparing its appearance from different angles, and if the object is made of glass or produces very bright reflections, the algorithms cannot compute an adequate 3D model from the images.

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Imagine a situation in which we are looking at a mirror object, i.e. one polished to such a degree that the roughness of its surface cannot be perceived by the eye. For example, a perfectly polished sphere (the simplest example is a Christmas-tree ornament in the form of a mirror ball). If there are no objects near such a sphere and we illuminate it with some light source in the dark, its outlines will be rather indefinite.

It should be noted that artists have often been interested in the problem of depicting mirror objects. Thus, images of objects of this kind appear in a number of paintings
by the well-known Russian artist K.S. Petrov-Vodkin, for example in the painting "With a Samovar", a reproduction of which is given here.

Why do we see a sphere, and not a luminous point at its center? If we place a perfectly polished sphere (of course, on a scale exceeding the resolving power of the eye) in the field of a plane wave (for example, by irradiating it with a laser beam) in a room with absorbing walls, we will indeed see nothing but a luminous point at the center of the sphere. The brightness of this point will be determined by the properties of the surface of the sphere, i.e. for example, a perfectly polished sphere of absorbing material would be perceived in the described situation as a dim point. In ordinary surroundings we see, for example, the same sphere surrounded by many other objects.
The objects depicted in it, whose images taken together form the visual image that we perceive as a sphere. If the surface of the sphere is rough, this essentially means that we are dealing with an object of nearly spherical shape with many singular points on its surface, which are sources of the scattered field.

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Generalizing the above reasoning to non-spherical objects allows us to understand in general terms why we see what we see. Indeed, for an object with a mirror surface the sources of the field scattered by it are located inside it. Therefore such an object is recognized the better, the more extraneous objects are near it and are reflected in it. Conversely, if an object has a rough surface, the sources of the scattered field lie directly on that surface, which makes the task of recognizing the shape of the object quite simple.

Thus, we come to the conclusion: for a scatterer to have the properties of an "invisible man", i.e. to be poorly recognizable, its shape must be as close as possible to an ideal analytic surface, with no corners, protrusions or other irregularities, and its surface must have absorbing
properties.

16. Problems of face recognition

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Researchers at Ben-Gurion University in Israel and the Japanese IT giant NEC found that carefully applied makeup on the forehead, cheeks and nose can help fool face recognition systems.

15.2 Images of the face of one person under different lighting conditions

Recognition Errors for Nested, Transparent, Mirrored and Other Objects

Ralph Gross, a researcher at the Robotics Institute of Carnegie Mellon, described in 2008 one obstacle related to the viewing angle of the face: " Face recognition becomes quite good on full frontal faces and 20 degrees, but once you go to profile, there were problems [ 8] ". In addition to variations in pose, low-resolution images are also very hard to recognize. This is one of the main obstacles to face recognition in surveillance systems .

Face recognition is less effective if there is facial expression. A big smile can make the system less effective. For example: Canada, in 2009, allowed only neutral facial expressions in passport photographs .

There is also inconsistency in the data sets used by researchers. Researchers may use anywhere from a few subjects to dozens of subjects and from a few hundred images to thousands of images. It is important for researchers to make the data sets they used available to one another, or to have at least a standard data set .

Data privacy is a major concern when it comes to storing biometric data in companies. Face data repositories or biometric data can be accessed by a third party if it is not stored properly or is broken into. In Techworld, Parris adds (2017), " Hackers will already be looking to reproduce people's faces in order to fool face recognition systems, but this technology has proved harder to break than fingerprint or voice recognition technologies have been in the past ".

Critics of the technology complain that the scheme in the London Borough of Newham, in operation since 2004, has never recognized a single criminal, despite the fact that several criminals in the system's database live in the borough and the system has been running for several years. "Not once, as far as the police know, has Newham's automatic face recognition system spotted a live target". This information seems to conflict with claims that the system was credited with a 34% reduction in crime (so why it was also rolled out in

продолжение следует...

Продолжение:


Часть 1 Recognition Errors for Nested, Transparent, Mirrored and Other Objects
Часть 2 - Recognition Errors for Nested, Transparent, Mirrored and Other Objects

See also

Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Pattern recognition"

Terms: Pattern recognition