Lecture
Internet search and retrieval systems (IPS) are specialized computer programs and services designed for searching and organizing information available on the Internet. They perform a key role in facilitating the searching, filtering and delivery of information to users in real time.
The Internet (Internet = inter + net – a union of networks) is a worldwide computer network that unites millions of computers into a single information system.
The Internet provides the widest possible opportunities for the free acquisition and dissemination of scientific, business and educational information (Example 13.1). The global network connects practically all
the major scientific and government organisations of the world, universities and business centres, information agencies and publishing houses, forming
a gigantic repository of data on all branches of human knowledge.
Virtual libraries, archives and news feeds contain a huge amount of text, graphic, audio and video information.
Example 13.1
The modern scientist and the Internet
«There are now 20—30 thousand journals published in the world, printing several million articles a year. One can get completely lost in this information flow, or one can feel quite at ease
by using electronic search systems. With the exception of a few eccentrics, no one goes to libraries anymore, no one leafs through the pages of journals,
no one works with cardboard bibliographic cards. All scientific journals have electronic counterparts, and a search engine picks out the articles you need in a matter of seconds. All that remains is to print them out and get on with the creative work – analysis. All this is done at your own computer, without getting up from your desk. And if you need a journal from times past, a manuscript or a section from a book that exists only in the
central library — even here there is no need to get up from your desk, let alone run or travel to the capital city. You find the central library's catalogue on the Internet, determine what you need,
and send a request by email. The librarian finds what is sought, scans it and sends it in electronic form straight to your computer. Fast, convenient and cheap. This is how most research workers in the world operate today». A. Demchenko .
According to experts' estimates, there are hundreds of millions of sites on the Internet, and this number doubles every year and a half.
So how does one find the information one needs in this gigantic repository of data? For this one needs to know how to use search systems.
A search engine is an online service that provides the ability to search for information on sites on the Internet.
Among search systems there are distinguished:
1) search engines in the pure sense (also called search machines, search engines);
2) classifiers (Internet catalogues, web directories, reference-type search tools);
3) metasearch systems.
Often one and the same portal contains both a search machine and a classifier.
A search machine is an online service that searches for information on the Internet by keywords and provides the user with a list of links to those sites that satisfy
the search criterion (Fig. 13.3). The main criteria of the quality of a search machine's operation are relevance (the correspondence of the result to the query), the completeness of the database, and allowance for language morphology.
A search machine performs a search for links in its own database, which is constantly updated – data on more and
more new sites is entered into it. The process of adding information about a site to a search engine's database is called indexing a site. The indexing of sites
is performed by a special program – a search robot (crawler).
A search robot (crawler) is a program that is a component part of a search system and is designed to traverse pages of the Internet in order to record information about them (keywords) in the search engine's database. The order in which pages are traversed, the frequency of visits, protection against looping, as well as the criteria for identifying keywords, are determined by the algorithms of the search machine. At present there exist several thousand search systems, however the majority of users turn to the services of approximately 10–15 of the most popular search engines:
here are some of the best-known Internet search and retrieval systems:
Google: Google is one of the most popular search engines in the world. It provides search results based on a range of factors, including the relevance and authority of websites.
Bing: Bing is a search engine developed by Microsoft. It provides search results and is also integrated into various Microsoft products, such as Windows and Office.
Yahoo: Yahoo Search provides search services and is one of the oldest players in this field. It also offers news, email and other services.
Yandex: Yandex is a Russian search engine that provides search and other online services. It is popular in Russian-speaking countries.
DuckDuckGo: DuckDuckGo is known for its focus on user privacy. It does not track or store users' personal information and provides anonymous search results.
Baidu: Baidu is the largest search engine in China. It provides search and other online services for Chinese users.
Wolfram Alpha: Wolfram Alpha provides knowledge-oriented results. It can perform calculations and provide structured answers to questions.
Startpage: Startpage provides anonymous searching of the Internet using Google's results, but without tracking users.
Ecosia: Ecosia is a search engine that promises to plant trees for every 45 search queries, in order to combat climate change.


A classifier is an online service that provides users with the addresses of sites and sometimes annotations of them, grouped into categories by subject. Each category may contain
several subcategories. By following the names of the headings, one can reach the information of interest. For example: Science – Economic Sciences – Management.
Classifiers can help a researcher in a case where they cannot precisely formulate a query but know the subject area in which they are searching for information (Fig. 13.4, 13.5)


A particular case of a classifier is the rating classifier.
A rating classifier is a classifier in which
the sites in the categories are sorted by degree of popularity, and
are also supplied with information about the number of visits to them (determined with
the help of visit counters).
A rating classifier makes it possible for the owners of their own pages, as well as for users, to quickly and precisely determine the number of visits to Internet pages. The rating classifier service
is provided, for example, on the Rambler portal (Fig. 13.6)

A metasearch system is an online service that makes it possible to search for information on the Internet using several search machines at the same time. A metasearch system
automatically forwards the query to real search machines and directories, and integrates the results received from them into a single whole.
As an example of a metasearch system one can name the MetaCrawler system (http://www.metacrawler.com).
A tool similar to a metasearch system is built directly into the interface of Internet Explorer. In order to gain access to this tool, you need to click the Search button on the Google Chrome toolbar. A «Search» panel will appear on the screen, containing a field intended for entering keywords. By entering a word or phrase in this field, you can see search results obtained using search engines such as Rambler, Yandex and Google.
These search and retrieval systems use various algorithms and methods to determine the relevance and ranking of search results. Users can use them to search for information on the Internet, as well as to gain access to a wide variety of online services and resources.
Comments