You get a bonus - 1 coin for daily activity. Now you have 1 coin

Computational Linguistics: Nature, Tasks, Branches and Applications

Lecture



Computational linguistics (also known as: mathematical or computational linguistics, English: computational linguistics) — a field of science concerned with the mathematical and computer modelling of intelligent processes in humans and animals in the creation of artificial intelligence systems, whose goal is to use mathematical models to describe natural languages.

Computational linguistics partially overlaps with natural language processing. However, the latter focuses not on abstract models but on applied methods of describing and processing language for computer systems.

The field of activity of computational linguists is the development of algorithms and applied programs for processing linguistic information.

Computational linguistics— a field of artificial intelligence whose goal is to use mathematical models to describe natural languages.

Computational linguistics partially overlaps with natural language processing. However, the latter focuses not on abstract models but on applied methods of describing and processing language for computer systems.

Broadly speaking, its field of activity is the development of algorithms and applied programs for processing linguistic information.

The task of computational linguists is to develop algorithms and applied programs for processing linguistic information.

Mathematical linguistics is a branch of artificial intelligence science. Its history began in the United States of America in the 1950s. With the invention of the transistor and the advent of a new generation of computers, as well as the first programming languages, experiments began with machine translation, especially of Russian scientific journals. In the 1960s, similar research was also carried out in the USSR (for example, an article on translation from Russian into Armenian in the 1964 collection «Problemy Kibernetiki» (Problems of Cybernetics)). However, the quality of machine translation still lags far behind translation produced by a human. By 2021, the quality of machine translation produced by the Google translator no longer lagged so far behind translation done by a human .

From 15 to 21 May 1958, the first All-Union Conference on Machine Translation was held at the 1st Moscow State Pedagogical Institute of Foreign Languages (1 MGPIIYa). The Organizing Committee was headed by V. Yu. Rozentsveig, with G. V. Chernov serving as the Committee's executive secretary. The full conference programme was published in the collection «Mashinnyi perevod i prikladnaya lingvistika» (Machine Translation and Applied Linguistics), issue 1, 1959 (also known as «Bulletin of the Machine Translation Association No. 8»). As V. Yu. Rozentsveig recalled, the published collection of conference abstracts reached the United States and made a great impression there.

In April 1959, the 1st All-Union Conference on Mathematical Linguistics was held in Leningrad, convened by Leningrad University and the Committee for Applied Linguistics. The Conference's main organiser was N. D. Andreev. A number of prominent mathematicians took part in the Conference, in particular S. L. Sobolev, L. V. Kantorovich (later a Nobel laureate), and A. A. Markov (the latter two spoke in the discussion). V. Yu. Rozentsveig gave the keynote address on the Conference's opening day, entitled «The General Linguistic Theory of Translation and Mathematical Linguistics».

Areas of computational linguistics

  • Natural language processing (NLP). Levels of text processing and analysis: syntactic, morphological, semantic.

The tasks and areas of computational linguistics include:

  1. Corpus linguistics — the creation and use of electronic text corpora.
  2. The creation of electronic dictionaries, thesauri, and ontologies. For example, Lingvo. Dictionaries are used, for example, for automatic translation and spell-checking.
  3. Automatic text translation. Among Russian translators, PROMT is popular. Among free translators, Google Translate is well known.
  4. Automatic fact extraction from text (information extraction) (English: fact extraction, text mining)
  5. Automatic summarisation (English: automatic text summarization). This function is included, for example, in Microsoft Word.
  6. Building knowledge management systems. See Expert systems.
  7. Creating question-answering systems (English: question answering systems).
  • Optical character recognition (English: OCR). For example, using the FineReader program
  • Automatic speech recognition (English: ASR).
  • Automatic speech synthesis.

Origins

Computational linguistics is often classed as part of artificial intelligence, but it existed before the advent of artificial intelligence. Computational linguistics originated in the United States in the 1950s with the aim of using computers for the automatic translation of texts from foreign languages, especially Russian scientific journals, into English. Since computers can perform arithmetic (systematic) calculations much faster and more accurately than humans, it was believed to be only a short matter of time before they would also be able to process language. [10]Computational and quantitative methods have also historically been used in attempts to reconstruct earlier forms of modern languages and to group modern languages into language families. Earlier methods, such as lexicostatistics and glottochronology, proved premature and inaccurate. However, recent interdisciplinary research, which borrows concepts from biological research, especially gene mapping, has yielded more sophisticated analytical tools and more reliable results. [11]

When machine translation (also known as mechanical translation) failed to immediately produce accurate translations, the automatic processing of human languages was recognised as far more complex than originally assumed. Computational linguistics was born as the name of a new field of research devoted to developing algorithms and software for the intelligent processing of linguistic data. The term «computational linguistics» itself was first introduced by David Hays, one of the founders of the Association for Computational Linguistics (ACL) and the International Committee on Computational Linguistics (ICCL). [12]

It has been observed that translating from one language to another requires understanding the grammar of both languages, including both morphology (the grammar of word forms) and syntax (the grammar of sentence structure). To understand syntax, it was also necessary to understand semantics and vocabulary (or the "lexicon"), and even something of the pragmatics of language use. Thus, what began as an attempt to translate between languages grew into an entire discipline devoted to understanding how to represent and process natural languages using computers. [13]

Today, research in computational linguistics is carried out in computational linguistics departments, [14] computational linguistics laboratories, [15] computer science departments, [16] and linguistics departments. [17] [18] Some research in computational linguistics is aimed at creating working speech- or text-processing systems, while other research aims to create a system that supports human–machine interaction. Programs designed for human–machine communication are called dialogue agents. [19]

Approaches

Just as computational linguistics can be pursued by specialists in a wide range of fields and departments, research areas can also span a broad spectrum of topics. The following sections discuss part of the literature available across the field, divided into four main areas of discourse: developmental linguistics, structural linguistics, linguistic production, and linguistic comprehension.

Developmental approaches

Language is a cognitive skill that develops throughout a person's life. This developmental process has been studied using several methods, and the computational approach is one of them. The development of human language imposes certain constraints that complicate the application of the computational method to its study. For example, during language acquisition, human children are mostly exposed only to positive data. [20] This means that during an individual's language development, only evidence of what constitutes a correct form is provided, with no evidence of what is incorrect. This information is insufficient for a straightforward hypothesis-testing procedure regarding information as complex as language [21] .and thus sets certain limits on the computational approach to modelling human language development and acquisition.

Attempts have been made to model the process of children's language-acquisition development from a computational perspective, leading to both statistical grammars and connectionist models. [22] Work in this area has also been proposed as a method for explaining the evolution of language over the course of history. Using models, it has been shown that languages can be learned through a combination of simple input data presented gradually as a child's memory improves and attention span increases. [23] This was at the same time seen as a reason for the extended developmental period of human children. [23] Both conclusions were drawn from the power of the artificial neural network that created the project.

Infants' ability to develop speech has also been modelled using robots [24] to test linguistic theories. With the ability to learn like children, the model was built on an affordance model, in which mappings between actions, perception, and effects were created and linked to spoken words. Importantly, these robots were able to acquire functioning mappings between words and meanings without needing grammatical structure, which greatly simplified the learning process and shed light on information that contributes to the modern understanding of language development. Importantly, this information could only be verified empirically using a computational approach.

As our understanding of human language development over the course of a lifetime continues to improve through the use of neural networks and learning robotic systems, it is also important to remember that languages themselves change and evolve over time. Computational approaches to understanding this phenomenon have yielded very interesting information. Using the Price equation and Polya urn dynamics, researchers have created a system that not only predicts future language evolution but also provides insight into the history of the evolution of modern languages. [25] This modelling effort achieved, through computational linguistics, what would otherwise have been impossible.

It is clear that our understanding of language development in humans, as well as over the course of evolutionary time, has improved fantastically thanks to advances in computational linguistics. The ability to model and modify systems at will gives science an ethical method for testing hypotheses that would otherwise be unresolvable.

Structural approaches

To build better computational models of language, it is crucial to understand the structure of language. To this end, English has been studied extensively using computational approaches in order to better understand how language works at the structural level. One of the most important aspects of studying linguistic structure is having access to large linguistic corpora or samples. This gives computational linguists the raw data needed to run their models and to better understand the underlying structures present in the vast amount of data contained in any single language. One of the most cited English linguistic corpora is the Penn Treebank. [26]Drawn from a wide variety of sources, such as IBM computer manuals and transcripts of telephone conversations, this corpus contains more than 4.5 million words of American English. This corpus has mainly been annotated using part-of-speech tags and syntactic bracketing, and this has yielded substantial empirical observations related to the structure of language. [27]

Theoretical approaches to the structure of languages have also been developed. This work gives computational linguistics a foundation for developing hypotheses that will advance the understanding of language in many ways. One of the original theoretical proposals concerning the interiorisation of grammar and language structure put forward two types of models. [21] In these models, learned rules or patterns are reinforced as their frequency of occurrence increases. [21] This work also raised a question for computational linguists to answer: how does an infant learn a specific, non-normalised grammar (Chomsky normal form) without learning an overgeneralised version and getting stuck? [21]Theoretical efforts like these set the direction for research at an early stage in the field's existence and are crucial to the field's growth.

Structural information about languages makes it possible to detect and perform similarity recognition between pairs of textual utterances. [28] For example, it has recently been demonstrated that, based on the structural information present in patterns of human discourse, recurrence conceptual graphs can be used to model and visualise trends in data and to create reliable measures of similarity between natural textual utterances. [28] This method is a powerful tool for further study of the structure of human discourse. Without a computational approach to this question, the extremely complex information present in discourse data would have remained inaccessible to researchers.

Information about the structural data of language is available for both English and other languages, such as Japanese. [29] Using computational methods, corpora of Japanese sentences were analysed, and a pattern of log-normality was found with respect to sentence length. [29]Although the exact cause of this log-normal distribution remains unknown, revealing this kind of information is precisely what computational linguistics is meant to do. This information could lead to further important discoveries regarding the basic structure of the Japanese language, and could have some bearing on the understanding of Japanese as a language. Computational linguistics makes it possible to add very interesting insights to the body of scientific knowledge quickly and reliably.

Without a computational approach to the structure of linguistic data, much of the information available today would still be hidden beneath the enormous volume of data within any given language. Computational linguistics allows scientists to analyse huge amounts of data reliably and efficiently, creating opportunities for discoveries that most other approaches cannot offer.

Production approaches

Producing language is just as complex, both in terms of the information it conveys and the skills a fluent speaker-producer must have. In other words, comprehension is only half of the communication problem. The other half is how the system produces language, and computational linguistics has made some interesting discoveries in this area.

Computational Linguistics: Nature, Tasks, Branches and Applications
Alan Turing: a computer scientist and namesake creator of the Turing test as a method of measuring machine intelligence.

In his now well-known paper, published in 1950, Alan Turing suggested that machines might one day be able to «think». As a thought experiment for determining what might define the concept of thinking in machines, he proposed the «imitation game», in which a person carries on two text-based conversations, one with another human and the other with a machine trying to respond like a human. Turing suggests that if the subject cannot distinguish the human from the machine, it can be concluded that the machine is capable of thought. [30] Today this test is known as the Turing test, and it remains an influential idea in the field of artificial intelligence.

Computational Linguistics: Nature, Tasks, Branches and Applications
Joseph Weizenbaum: a former professor at the Massachusetts Institute of Technology and computer scientist who developed ELIZA, a primitive computer program that used natural language processing.

One of the earliest and best-known examples of a computer program designed for natural communication with people is the ELIZA program, developed by Joseph Weizenbaum at the Massachusetts Institute of Technology in 1966. The program imitated a Rogerian psychotherapist when responding to the user's written statements and questions. It seemed to be able to understand what it was being told and respond sensibly, but in reality it simply followed a pattern-matching procedure based on recognising only a few key words in each sentence. Its responses were generated by recombining the unrecognised parts of a sentence around correctly translated versions of the recognised words. For example, in the phrase «You seem to hate me», ELIZA recognises «you» and «me», which matches the general pattern «you [some words] me», allowing ELIZA to swap «you» and «me» for «I» and «you» and reply: «Why do you think I hate you?». In this example, ELIZA does not understand the word «hate».[31]

Some projects are still trying to solve the problem that first made computational linguistics its own field. However, the methods have become more sophisticated, and as a result the findings produced by computational linguists have become more informative. To improve machine translation, a comparison was made of several models, including hidden Markov models, smoothing methods, and their specific refinements for application to verb translation. [32] The model found to produce the most natural translations from German and French was an improved alignment model with first-order dependency and a fertility model. They also provide efficient training algorithms for the models presented, which can give other researchers the means to improve their own results. This type of work is characteristic of computational linguistics and has applications that can significantly improve the understanding of how language is produced and understood by computers.

Work has also been done aimed at getting computers to produce language in a more natural manner. Using linguistic input from people, algorithms were created that can change the system's production style based on a factor such as human linguistic input, or on more abstract factors such as politeness or any of the five main personality traits. [33] This work uses a computational approach based on parameter-estimation models to classify the wide range of linguistic styles we observe in different people and to simplify it for computer processing in the same way, making human-computer interaction far more natural.

Text-based interactive approach

Many of the earliest and simplest models of human-computer interaction, such as ELIZA, for example, involve the user typing text to receive a response from the computer. With this method, the words entered by the user cause the computer to recognise certain patterns and respond accordingly through a process known as keyword detection.

Speech-based interactive approach

The latest technologies place greater emphasis on speech-based interactive systems. These systems, such as Siri on the iOS operating system, work on the same pattern-recognition method as text-based systems, but in the former the user input is carried out through speech recognition. This branch of linguistics involves processing the user's speech as sound waves and interpreting acoustic and linguistic patterns so that the computer can recognise the input. [34]

Approaches to understanding

Much of the attention of modern computational linguistics is focused on understanding. With the spread of the Internet and the abundance of readily available written text, the ability to create a program capable of understanding human language would open up many broad and exciting possibilities, including improved search engines, automated customer service, and online learning.

Early work on understanding involved applying Bayesian statistics to the task of optical character recognition, as demonstrated by Bledsoe and Browning in 1959, when a large dictionary of possible letters was generated by «training» on example letters, after which the probabilities that any of these learned examples matched new input data were combined to reach a final decision. [35] Other attempts to apply Bayesian statistics to language analysis included the work of Mosteller and Wallace (1963), in which an analysis of the words used in «The Federalist Papers» was used to try to determine their authorship (concluding that Madison was most likely the author of the majority of the essays). [36]

In 1971, Terry Winograd developed an early natural language processing engine capable of interpreting naturally written commands within a simple, rule-governed environment. The core language-parsing program in this project was called SHRDLU and was able to hold a fairly natural dialogue with a user giving it commands, but only within a toy environment designed for this task. This environment consisted of blocks of various shapes and colours, and SHRDLU could interpret commands such as «Find the block that is taller than the one you are holding and put it in the box», and ask questions such as «I don't understand which pyramid you mean» in response to the user's input. [37]Impressive as this kind of natural language processing is, it proved to be far more difficult outside a toy environment. Similarly, a NASA project called LUNAR was developed to give answers to naturally phrased questions about the geological analysis of lunar rocks returned by the Apollo missions. [38] This kind of task is called question answering.

Initial attempts to understand spoken language were based on work carried out in the 1960s and 1970s on signal modelling, in which an unknown signal is analysed to find patterns and make predictions based on its history. An initial and somewhat successful approach to applying this kind of signal modelling to language was achieved using hidden Markov models, described in detail by Rabiner in 1989. [39] This approach attempts to determine the probabilities for an arbitrary number of models that could be used to generate speech, as well as to model the probabilities of the various words generated from each of these possible models. Similar approaches were used in early attempts at speech recognition, starting in the late 1970s at IBM, using the probabilities of word/part-of-speech pairs.[40]

More recently, these kinds of statistical approaches have been applied to more complex tasks, such as topic identification using Bayesian parameter estimation to infer topic probabilities in text documents. [41]

Applications

Applied computational linguistics is largely equivalent to natural language processing. Examples of end-user applications include speech recognition software such as Apple's Siri feature, spell-checking tools, speech synthesis programs that are often used to demonstrate pronunciation or to assist people with disabilities, and machine translation programs and websites such as Google Translate. [42]

Computational linguistics is also useful in situations involving social media and the Internet, for example in providing content filters for chats or website search, [42] in grouping and organising content through data mining in social networks, [43] in document retrieval and clustering. For example, if a person searches for «red, big, four-wheeled car» to find pictures of a red truck, the search engine will still find the relevant information by matching words such as «four-wheeled» with «car». [44]

Computational approaches are also important for supporting linguistic research, for example in corpus linguistics or historical linguistics. As for the study of change over time, computational methods can help with modelling and identifying language families (see further quantitative comparative linguistics or phylogenetics), as well as modelling changes in sound [45] and meaning. [46]

Major associations and conferences

  • The Association for Computational Linguistics (ACL): divided into two branches, European and North American .
  • The «Dialogue» International Conference on Computational Linguistics .
  • The International Conference on Computational Linguistics and Intelligent Text Processing[en] (CICLing).

See also

  • Applied linguistics
  • Corpus linguistics
  • Text theory
created: 2022-01-18
updated: 2026-03-10
132



Was this answer useful?
Choose a quick rating so we can improve the next answer for you.
How satisfied are you?


Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Computational linguistics"

Terms: Computational linguistics