In linguistics, a corpus (in this meaning, the plural is stressed as кóрпусы, not корпусá) — a body of texts compiled and processed according to specific rules, used as a basis for studying a language. Corpora are used for statistical analysis, for testing statistical hypotheses, and for confirming the linguistic rules of a given language. The text corpus is the subject of study of corpus linguistics.
Basic properties of a corpus
Among the many definitions of a corpus, its main properties can be identified:
- electronic — in the modern sense a corpus must be in electronic form
- representative — it must accurately «represent» the object it models
- annotated — the main distinction between a corpus and a plain collection of texts
- pragmatically oriented — it must be created for a specific task
Classification of corpora
Corpora can be classified according to various criteria: the purpose for which the corpus was created, the type of language data, «literariness», genre, dynamism, type of annotation, volume of texts, and so on. By the criterion of parallelism, for example, corpora can be divided into monolingual, bilingual and multilingual ones. Multilingual and bilingual corpora are divided into two types:
- parallel — a set of texts and their translations into one or more languages
- comparable (pseudo-parallel) — original texts in two or more languages
Corpus annotation
Annotation consists in attaching special tags to texts and their components: linguistic and extralinguistic (external) ones. The following linguistic types of annotation are distinguished: morphological, semantic, syntactic, anaphoric, prosodic, discourse, and so on. Further structural levels of analysis are applied to some corpora. In particular, some small corpora may be fully syntactically annotated. Such corpora are usually called deeply annotated or syntactic, and the syntactic structure itself in this case is a dependency tree.
Manual annotation (tagging) of texts is an expensive and labor-intensive task. At present, various software tools for annotating corpora are available in the public domain . They can conventionally be divided into stand-alone and web-based. In recent years, the focus of developers has shifted toward web applications. These systems offer a number of advantages:
- the possibility of several people annotating the same document simultaneously
- no need to install additional software besides a browser
- flexible access rights management
- display of the current progress of the annotation process
- the possibility of modifying the corpus being annotated
The internet as a corpus
Modern technologies make it possible to create «web corpora», that is, corpora obtained by processing internet sources:
A web corpus is a special kind of linguistic corpus created by progressively downloading texts from the internet using automated procedures that determine the language and encoding of individual web pages on the fly, remove templates, navigation elements, links and advertising (so-called boilerplate), and perform text transformation, filtering, normalization and deduplication of the resulting documents, which can then be processed with traditional corpus-linguistics tools (tokenization, morphosyntactic and syntactic annotation) and loaded into a corpus search system. Building a web corpus is not only much cheaper, but above all its size can be an order of magnitude larger than that of traditional corpora .
— Vladimír Benko ARANEA — A FAMILY OF BILLION-WORD WEB CORPORA
Speech corpus
Speech corpus (sound corpus) — a database of audio files and text transcriptions, a variety of text corpus. In speech technologies[en] speech corpora are used, among other things, to build acoustic models[en] (which can then be used in speech-recognition engines). In linguistics, speech corpora are used for research in phonetics, dialectology, conversation analysis and other fields.
There are two types of speech corpora:
1.Databases of read-aloud texts, including:
- book texts;
- news broadcast texts;
- word lists;
- number sequences.
2.Databases of recordings of spontaneous speech — including:
- dialogues — conversations between two or more people;
- oral narratives (for example, Buckeye Corpus );
- map-task explanations — one person explains a route on a map to others;
- appointment tasks — two people try to find a common meeting time based on separate schedules.
A particular kind of speech corpus consists of databases of texts spoken by people who are not native speakers of the language[en], which contain speech with a foreign accent.
Applications
A corpus is the basic concept and database of corpus linguistics. The analysis and processing of various types of corpora are the subject of most work in the field of computational linguistics (for example, keyword extraction), speech recognition and machine translation, in which corpora are often used to build hidden Markov models for part-of-speech tagging and other tasks. Corpora and frequency dictionaries can be useful in foreign-language teaching.
Corpora are the primary knowledge base in corpus linguistics . Other well-known areas of application include:
- Language technologies , natural language processing , computational linguistics
- The analysis and processing of various types of corpora is also the subject of extensive work in computational linguistics , speech recognition and machine translation , where they are often used to build hidden Markov models for part-of-speech tagging and other purposes. Corpora and the frequency lists derived from them are useful for language teaching . Corpora can be regarded as a type of aid for writing in a foreign language, since the contextualized grammatical knowledge that non-native users acquire through exposure to authentic texts in corpora allows learners to understand how sentences are formed in the target language, enabling effective writing.
- Machine translation
- Multilingual corpora specially formatted for parallel comparison are called aligned parallel corpora . There are two main types of parallel corpora containing texts in two languages. In a translation corpus texts in one language are translations of texts in another language. In a comparable corpus texts are of the same kind and cover the same content, but they are not translations of one another. To use parallel text, a prerequisite for analysis is some form of text alignment that identifies equivalent text segments (phrases or sentences). Machine translationAlgorithms for translating between two languages are often trained using parallel fragments comprising a corpus in the first language and a corpus in the second language, which is an element-by-element translation of the first-language corpus.
- Philology
- Text corpora are also used in the study of historical documents , for example in attempts to decipher ancient scripts or in biblical studies . Some archaeological corpora may be so short that they give only a snapshot in time. One of the shortest corpora in terms of time span may be the Amarna letters, covering 15–30 years ( around 1350 BC ). The corpus of an ancient city, (for example, the « Kültepe texts» from Turkey), may pass through a series of corpora, determined by the date the site was found.
English
- American National Corpus
- Bank of English
- British National Corpus
- Bergen Corpus of London Teenage Language (COLT)
- Brown Corpus , part of the "Brown Family" of corpora, together with LOB , Frown and F-LOB
- Corpus of Contemporary American English (COCA), 425 million words, 1990–2011. Free search online
- Corpus Resource Database (CoRD), more than 80 English-language corpora.
- GUM Corpus , a multilayer open-source corpus from Georgetown University, with a very large number of annotation layers
- Google Books Ngram Corpus
- International Corpus of English
- Oxford English Corpus
- RE3D (Relationship and Entity Extraction Evaluation Dataset)
- Santa Barbara Corpus of Spoken American English
- Scottish Corpus of Texts and Speech
European languages
- CETENFolha
- Corpus of Electronic Texts
- Corpus Inscriptionum Insularum Celticarum (CIIC), covering primitive Irish ogham inscriptions
- Google Books Ngram Corpus
- Corpus of the Georgian Language
- Thesaurus Linguae Graecae (Ancient Greek)
- Eastern Armenian National Corpus (EANC) 110 million words. Free search online.
- Spanish Text Corpus from Molino de Ideas, containing 660 million words.
- CorALit: Corpus of Academic Lithuanian Texts, published in 1999–2009 (about 9 million words). Compiled at Vilnius University, Lithuania
- Reference Corpus of Contemporary Portuguese (CRPC)
- Turkish National Corpus
- CoRoLa - Reference Corpus of Contemporary Romanian (Corpus Representzentativ al limbii române contemporane)
- TS Corpus - a large collection of Turkish corpora. TS Corpus is a free and independent project whose goal is to create Turkish corpora, NLP tools and linguistic datasets ...
Slavic
East Slavic
- Belarusian N-Corpus
- Russian National Corpus
- General Internet Corpus of the Russian Language
- General Regionally Annotated Corpus of the Ukrainian Language
- Corpus of the Ukrainian Language
- Araneum Russicum
- Russian Corpus of Biographical Texts
- RuTweetCorp
- RusAge: Corpus for Age-Based Text Classification
- National Corpus of the Russian Language
- General Internet Corpus of the Russian Language
- Russian-Language Corpus of the Aranea Project
- Corpus of Biographical Texts
- RuTweetCorp
South Slavic
- Bulgarian National Corpus
- Corpus of the Croatian Language
- Croatian National Corpus
- Slovenian National Corpus
West Slavic
- Czech National Corpus [10]
- National Corpus of Polish
German
- German Reference Corpus (DeReKo), more than 4 billion words of contemporary written German.
- Free corpus of German errors by people with dyslexia
Middle Eastern languages
- Corpus Inscriptionum Semiticarum
- Kanaanäische und Aramäische Inschriften
- Hamshahri Corpus ( Persian )
- Persian in the MULTEXT-EAST corpus (Persian) [11]
- Amarna letters (for Akkadian , Egyptian, Sumerograms, etc.)
- TEP: Tehran English-Persian Parallel Corpus [12]
- TMC: Tehran Monolingual Corpus , a standard corpus for Persian language modeling [12]
- Persian Today Corpus: the most frequent words in Persian today, based on a corpus of one million words (in Persian: Vāže-hā-ye Porkārbord-e Fārsi-ye Emrūz ), Hamid Hassani , Tehran, Iranian Language Institute (ILI) , 2005, 322 p. ISBN 964-8699-32-1
- Kurdish-corpus.uok.ac.ir (Kurdish corpus, Sorani dialect) University of Kurdistan, Faculty of English Language and Linguistics
- Bijankhan Corpus A Contemporary Persian Corpus for NLP research, University of Tehran , 2012
- Open Neo-Assyrian Texts Corpus Project
- Quranic Arabic Corpus (Classical Arabic)
- Electronic Text Corpus of Sumerian Literature
- Open Richly Annotated Cuneiform Corpus
- Asosoft Text Corpus [13]
Devanagari
- Nepali Text Corpus (90+ million running words / 6.5+ million sentences)
East Asian languages
- Kotonoha Japanese Language Corpus [14]
- LIVAC Synchronous Corpus (Chinese)
South Asian languages
- SinMin Dataset [15] ( Sinhalese )
See also
- Computational linguistics
- Keyword
- Concordance
- Corpus linguistics
- Linguistic Data Consortium
- Natural language processing
- Natural language toolkit
- Parallel text alignment
- Speech corpus
- Translation memory
- Treebank
- Zipf's law
- Arabic Speech Corpus
- Common Voice
- EXMARaLDA
- List of children's speech corpora
- Non-native speech database
- Praat
- Spoken English Corpus
- BABEL Speech
- TIMIT
- Transcriber
- Transcription (linguistics)
Comments