An ontology of data analysis

Lecture



"Why not pull it out? It can be pulled out. Only here one must understand — it cannot be done without some notion of the matter… Teeth come in different kinds. One you pull with forceps, another with a goat's foot, a third with a key… Depends which."
A. P. Chekhov

Introduction

Streams of textual and numerical information are generated daily and settle into data warehouses. How fully, in practice, are all the patterns hidden within this data – patterns that may well be of great value – actually used? One might assume that the share of "raw" data converted into practically useful knowledge is still quite modest. Even the rich arsenal of classical statistics is far from fully used, to say nothing of more modern methods of nonlinear analysis. "Where one is obliged to worship the sun, the laws of heat will be but poorly understood." The point is that in our country, although statistics was never branded a "venal wench of the bourgeoisie", for a long time there was a rejection of formal statistics. What statistics could there be, when the data themselves had to conform to the state's ideological guidelines. The situation is compounded by the fact that new methods of data analysis and knowledge extraction have recently been developing actively, based on approaches other than the traditional integro–differential paradigm. These are the methods of evolutionary modelling and machine learning. The term "evolutionary modelling" is now fairly well established, and it is generally understood to mean genetic algorithms and artificial neural networks. The term "machine learning" leaves more room for debate about which methods it covers; in particular, decision trees belong here.

What is an ontology?

How does one find one's bearings amid this variety of tools? Which of them should be chosen to solve a particular problem? In this situation the comparatively new term – "ontology" – comes in very handy. An ontology is a precise specification of a given domain. It provides a vocabulary for representing and exchanging knowledge about that domain, together with a set of relations established between the terms in that vocabulary. In the simplest case, building an ontology comes down to:

  • Identifying concepts – the basic notions of a given domain;
  • Establishing connections between concepts – defining the relations and interactions of the basic notions.

One of the advantages of using ontologies as a tool of cognition is a systemic approach to studying the domain. This achieves:

  • Systematicity – an ontology presents a coherent view of the domain;
  • Uniformity – material presented in a single form is understood and reproduced far better;
  • Rigour – building an ontology makes it possible to restore the missing logical connections in their entirety.

What is data analysis?

Data analysis — a field of mathematics and computer science concerned with the construction and study of the most general mathematical methods and computational algorithms for extracting knowledge from experimental (in the broad sense) data; the process of investigating, filtering, transforming and modelling data with a view to extracting useful information and making decisions.

Put a little more simply, I would propose to understand data analysis as the set of methods and applications associated with algorithms for processing data that do not have a strictly fixed answer for every incoming object. This is what distinguishes them from classical algorithms, for example those implementing sorting or a dictionary. The execution time and memory footprint of a classical algorithm depend on its particular implementation, but the expected result of applying it is strictly fixed. By contrast, we expect a neural network recognising digits to answer 8 when shown a handwritten eight, but we cannot require this result. Moreover, any (in a reasonable sense of the word) neural network will occasionally err on some correct variants of input data or other. Let us call this formulation of the problem, and the methods and algorithms applied in solving it, non-deterministic (or fuzzy) as opposed to classical (deterministic, crisp) ones.

An ontology of data analysis

Since knowledge is personal in character, one and the same domain can be described by different ontologies. This is especially true of poorly formalisable domains, or where there are a great many contentious issues.

An ontology of data analysis

Mathematical statistics

To solve problems involving the analysis of data in the presence of random and unpredictable influences, mathematicians and other researchers have, over the past two hundred years, developed a powerful and flexible arsenal of methods known collectively as mathematical statistics. Over this time a great deal of experience has accumulated in the successful application of these methods in various spheres of human activity, from economics to space research. And under certain conditions these methods make it possible to obtain optimal solutions. For example, one of the problems solved in radar – detecting a known signal against a background of additive interference in the form of white noise. Methods of mathematical statistics solve this problem in an optimal way, and it is difficult to imagine the need for other approaches to solving it. At the same time, the problem of resolving closely spaced targets under more complex interference conditions is solved less successfully by linear statistical methods.

Evolutionary modelling

Nowadays, when speaking of evolutionary modelling, one usually means genetic algorithms and artificial neural networks. The term "evolutionary modelling" owes its origin to the source from which the ideas underlying this paradigm were borrowed. Whereas classical approaches are based on a human's knowledge of the domain, formalised in some way, for a neural network an analytical form of representing knowledge is unavailable; all it can do is memorise and generalise the empirical dependencies between input factors and resulting values presented to it during training. That is, a neural network builds a model of some process and subsequently reproduces its behaviour. This gives some researchers grounds to claim that artificial neural networks model modes of thought characteristic of humans. In our view, for the practical use of neural-network technologies it is enough that neural networks are capable of building complex nonlinear models of processes, while how human brains are actually arranged is a secondary matter. What matters is something else – the quality of a model depends on the quality of the training data (here it is much the same as with people).

Genetic algorithms use the mechanisms of genetic evolution, which can in general be formulated as follows: the higher an individual's fitness, the higher the probability that in its offspring this fitness will be expressed even more strongly. Treating the process of adaptation as an optimisation process leads to the idea of using genetic algorithms when training neural networks. Moreover, whereas gradient training methods guarantee finding a local minimum, a genetic algorithm provides global optimisation.

Field of application

Methods of evolutionary modelling solve a wide class of problems: pattern classification, clustering, approximation, data forecasting, optimisation, associative memory, control of dynamic objects. Moreover, for all the reasons given above, neural networks cope with the tasks listed more successfully than methods of mathematical statistics do, the less formalisable the task is.

Advantages of neural networks

  • One of the main advantages of neural networks is that they have a broad field of application. Decision trees, by contrast, are confined within the bounds of classification tasks; it should be noted that algorithms exist that solve forecasting problems, but they fall considerably short of neural networks;
  • By their nature neural networks are universal approximators and make it possible to model very complex patterns, something that, say, classical regression models cannot achieve;
  • There is no need to know the form of the function being approximated in advance;
  • A neural network can easily be retrained to take newly arrived data into account; for decision trees this is, to this day, a major problem, since no method has been developed for "extending" a tree – one has to build the tree from scratch, without taking the previously built one into account;
  • There exist neural-network paradigms, for example Kohonen maps, in which the training process takes place without a teacher, i.e. the network works out the structure of the data by itself;
  • Another neural-network paradigm, RBF networks, trains very quickly, although it should be noted that the so-called "curse of dimensionality" affects them to a greater degree.

Machine learning

The goal of machine learning methods is to obtain simple classifying expressions that would be easy for a human to understand. The advantage of such methods is that human involvement is not required while the method is running.

Field of application

In a study conducted within the European StatLog project, an analysis was carried out of statistical methods (discriminant analysis, cluster analysis, etc.), decision trees (C4.5, AC2, CART, NewID, CN2, Itrule, etc.) and neural networks (multilayer networks, RBF networks, Kohonen maps) for solving classification problems. The data were taken from various domains: pattern recognition (handwritten text, vehicles), medical diagnosis (diabetes, head injuries, heart disease), molecular biology (recognising DNA structure), credit granting, and so on.

The study found that decision trees produced the best results in solving the following problems:

  1. Assessing the creditworthiness of a candidate for a loan;
  2. Diagnosing faults in technical systems;
  3. Placing radiators on the Space Shuttle.

Advantages of decision trees

  • Training decision trees takes far less time than, for example, training neural networks;
  • The result is presented in a form that is easy for a human to interpret. A classification model presented as a tree is intuitively understandable to a person, unlike neural networks, which are by their nature a black box;
  • Any number of parameters can be fed to the input of a decision-tree algorithm; the algorithm itself will select the most significant parameters, and only these will appear in the resulting tree. This relieves the user of the need to determine the input parameters. Here again, when using neural networks we must approach the question of input fields very carefully, since as the number of input fields grows, the time spent on the training process increases – and that process is already very long and gives rise to many complaints;
  • The forecast accuracy of decision trees is comparable to that of other methods of building classification models (statistical methods, neural networks);
  • There exist scalable decision-tree algorithms, SLIQ and SPRINT, i.e. as the number of examples grows the time spent on training grows linearly, for building decision trees on very large databases;
  • Decision-tree building algorithms have methods for the special handling of missing data;
  • Classical and modern statistical methods used in classification tasks work only with numerical data; decision trees work successfully with both numerical and string values. Moreover, some statistical methods are parametric, i.e. we must know in advance the form of the model, or the dependency between the dependent and independent variables. For example, classifiers built on the maximum-likelihood principle assume that the data have a normal distribution;
  • They make it possible to extract rules in natural language, for example: If age > 35 AND income > Average Then Grant the loan.

Conclusion

On our forum one can sometimes come across rather irritated remarks about all these clevernesses. For some reason neural networks are held in particular dislike. We would like to call on these commenters to show a little more restraint, and to say the following.

First, if one looks around soberly, it turns out that a few magic words, such as neural network, perceptron, factor analysis, regression analysis… , cannot solve all unsolved problems. "It is very rarely that one succeeds in unlocking several of nature's secrets at once with the same key." (C. Shannon).

Second, the effectiveness of nonlinear estimation techniques (meaning neurocomputing) can be increased by combining them with already familiar linear statistical methods. An example – RBF networks, in which the hidden-layer weights are tuned with the help of a genetic algorithm, while the output-layer weights are computed by the good old method of pseudo-inverse matrices.

It is merely a tool. How to use it is, in the end, for the person to decide. Incidentally, the story told by Chekhov in his tale "Surgery" (from which the epigraph is taken) happened only because, instead of the doctor, who had gone off to get married, patients were being seen by the feldsher Kuryatin.

Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Data mining"

Terms: Data mining