Lecture
A model in linguistics is a logical (sign) construct (device), artificially created for linguistic purposes, which reproduces certain characteristics of the object under study, subject to predetermined requirements for the correspondence of this construct to the object.
Despite the apparent complexity of this definition, we can easily create models of objects, sometimes without even realizing that we are doing so – though not at the scientific level, but at the practical, everyday one. For example, when explaining how to get somewhere, we might draw a plan, which is essentially a model of the real terrain. It does not reproduce the terrain in full, but depicts only those elements that are significant for solving the task at hand – in this case, finding the way somewhere. The most important thing is that a model must necessarily correspond to the purpose and conditions of the modelling.
Every model is built on the basis of a hypothesis about the possible structure of the original and represents a functional analogue, a mock-up of the original. This makes it possible, having studied the model, to transfer knowledge from the model to the original. A practical experiment can serve as a criterion for the adequacy of a model.
Modelling an object is a necessary element of understanding it. But this is only one of the stages of cognition. An infinite number of models of one and the same object can be constructed, all corresponding to its properties yet differing from one another, since a model reflects not only the properties of the object but also our point of view on it. Moreover, a model must satisfy certain requirements regarding its correspondence to the object being modelled.
For example, a written (printed) text recording oral speech, and a phonetic transcription of oral speech, are sign models of one and the same real event, but the requirements for their correspondence to the object differ.
No model is complete. A complete description is neither possible nor necessary, since in modelling we single out only certain properties of the object that are essential to us. Modelling the same properties, even within one and the same science, can lead to different models; the differences between them depend on the objectives of the modelling.
For example, the model of the phoneme system of the Russian language differs between the “Leningrad” and “Moscow” phonological schools.
In science, the fundamental principle of the multiplicity of models of one and the same object is becoming ever more firmly established. It should also be understood that every model is optimal for a specific purpose: if, for example, a theoretical grammar were built into a computer program, it would prove useless.
By artificial intelligence should be understood attempts to build working models – that is, models implemented on a computer – of various aspects of human linguistic activity. This area is the subject matter of computational linguistics, which as a whole studies linguistic phenomena by the method of computer modelling.
In its applied aspect, computational linguistics sets itself the following tasks:
The creation of natural language processing systems (in written and spoken form);
The development of linguistic technologies and resources (above all, dictionaries and text corpora);
Computer support for linguistic research.
Modelling of natural language covers, first and foremost, the analysis (comprehension) and synthesis (production) of texts. Both procedures are carried out in several stages. The first can be modelled as follows:
a text is received at the input. It undergoes morphological analysis (for example, the noun, its case, gender, and number are determined);
syntactic analysis of the text: the syntactic structures of sentences are constructed;
semantic analysis: the meanings of words are identified;
pragmatic processing of the text (taking contextual information into account).
Production models consist of the same stages, but solve the inverse tasks: the pragmatic context determines the plan of the future text and generates semantic representations of individual words, linking them together; the syntactic component arranges the words in the required order, and the morphological component builds the specific word forms.
The quality of a language model depends on the depth of the linguistic theory underlying it and on the formal language in which the linguistic information (the model’s grammar) is recorded.
Various applied natural language processing systems exist. They include spelling- and grammar-correction modules connected to word processors, machine translation systems, systems for planning and generating text from a non-linguistic source representation (for example, from graphics), information extraction and automatic summarization systems, question-answering systems, natural-language interfaces to databases, speech recognition and synthesis, and others.
Word processors, referred to as automatic text-processing systems (Eng. text processing or word processing systems), form part of the system software of practically all types of computers. The central function of word processors is the processing of text by specialized programs (as described above). Such programs use, as reference data, the formal grammars and dictionaries of the language whose texts serve as the object of analysis or synthesis.
Interlinguistics is a special linguistic discipline concerned with organizing effective communication in a multilingual world and with the prospects for the linguistic development of humanity. The area of scholarly interest of interlinguistics is the elaboration of principles and methods for creating artificial languages.
The idea of creating a universal artificial language arose long ago, from the time humanity became aware of the multilingualism of the world and encountered the practical inconveniences associated with it. Creating a single language is a task that is not only scientific and theoretical but also social. Throughout the whole of human history, the idea of creating such a language has repeatedly arisen, justified by the practical needs of society. However, the languages created in the Middle Ages, and in the seventeenth and eighteenth centuries, remained mere experiments. In the nineteenth century, as capitalism developed, ties between states grew stronger, and people increasingly felt a practical need for the creation of a single artificial language that would simplify international communication.
In 1879–1880, Johann Schleicher created the artificial language Volapük by blending and truncating words taken from the Romance and Germanic languages, and by radically simplifying the grammar. This was the first artificial language to achieve fairly wide currency: it was mastered by about 210,000 people, 30 newspapers and magazines were published in it, and 300 literary works were translated into it. Despite its seemingly simple appearance, Volapük was difficult as an artificial language. It lacked a sound scientific basis for word formation.
In 1887, the Warsaw ophthalmologist Ludwik Zamenhof invented a new international language, Esperanto. This was a successful attempt at creating an artificial language. Esperanto is grammatically simple (16 rules) and logical in its word formation. Parts of speech are determined by the ending added to the stem. Verbs have only three simple tenses. Esperanto won wide popularity. Many great figures studied and used Esperanto: Jules Verne, Albert Einstein, Maxim Gorky. The largest Esperantist library in London holds tens of thousands of titles. In 1987 the 100th anniversary of Esperanto was celebrated. Despite the fact that a great many people around the world know this language, it cannot be said that Esperanto has truly become a universal language. Of course, it does solve the practical problem of interlingual communication, but, unfortunately, only within limited spheres. As a rule, Esperanto is used by specialist collectors (philatelists, numismatists, and others) or by those who pursue it simply “for the love of it”.
Esperanto did not remain the only artificial language on Earth. In 1907 it was reformed in the area of word formation, and as a result the language Ido appeared. At present, more than 500 projects for new artificial languages are known.
There is a great deal of debate – for and against – the creation of artificial languages. Supporters of an artificial language point to its extreme logicality, the simplicity of its grammar, the internal consistency of its lexical system (the absence of synonyms and homonyms), the absence of difficulties in mastering it, and certain other advantages.
Opponents of an artificial language argue quite convincingly for its drawbacks: artificial languages lack the idiomatic quality that accounts for the uniqueness of every natural language. Moreover, adopting an artificial language would mean a break with the whole of world culture, and the transition to an artificial language would require enormous material costs (translating a huge mass of information, retraining teachers, and so on).
The solution to the main problem of interlinguistics (the creation of a single artificial language) still remains hypothetical. For this reason, two types of languages are currently discussed:
Languages with maximally broad functions, of the Esperanto type;
Languages with maximally narrow functions, devoid of a sound form – abstract machine languages, intermediary languages used in machine translation, and algorithmic programming languages for computers.
In the near future, it is languages of precisely this type whose use is most realistic.
The word statistics immediately brings computation to mind. Indeed, linguo-statistical methods are used in linguistics for processing and analysing linguistic material, the quantity of which can be very large. Mathematical linguistics is able to describe and represent linguistic facts effectively thanks to such properties of language as its systemic character and the generalized nature of its units. Mathematical linguistics draws on various branches of mathematics: algebra, set theory, logic, mathematical statistics, and certain others.
Linguo-statistical methods are used in linguistics to reveal the connection between the quantitative and qualitative aspects of language: between the frequency of use and the age of words, and between the length of a word and its frequency of use. Phonetics and the morphemic composition of the word are studied statistically (for example, establishing the connection between the number of phonemes and the average length of morphemes). Rhythm and rhyme likewise possess quantitative parameters.
Quantitative methods are most widely used in describing the lexical level of the language system. The practical result of the statistical study of vocabulary is the frequency dictionary. Quantitative methods have become more effective with the advent of computing technology. Mathematical linguistics should not be regarded as a separate field of knowledge, as a separate linguistics. Linguistics simply uses mathematical methods and applies them to language in order to represent, by convenient means, those relationships for which ordinary verbal expressions would prove imprecise or overly complex.
In its use of mathematical methods, linguistics keeps pace with the times: whereas in the past the simplest statistical procedures were used to study language, linguists today have at their disposal a huge number of computer programs capable of solving the most diverse linguistic tasks. Some of these programs allow standard data processing to be carried out, similar to that usually performed in other sciences as well. But there are also specially developed programs that solve specific linguistic tasks. Thanks to modern technology, present-day linguistics can more effectively solve not only theoretical but also practical tasks. In particular, linguists now have new capabilities in preparing and compiling dictionaries, creating text corpora, and so on.
In the assignment for one of the previous lectures, you were asked to represent graphically the system and the structure of language – that is, to construct a model of language. Which essential properties of the object (language) are singled out for constructing this model?
How do you understand the goals of creating artificial intelligence? What tasks are the researchers working in this field trying to solve?
Do you think humanity needs a single language, and can one be created? If it is possible, would people want to use it?
If you were to undertake the creation of an artificial language, what principles would you build it on?
Which properties and aspects of language do you think could be described mathematically?
Comments