Information and data processing

Lecture



Information processing — the entire set of operations (collection, input, recording, transformation, retrieval, storage, destruction, registration) carried out with the help of technical and software tools, including exchange over data-transmission channels.

Data processing includes many types depending on the goals and methods of working with information. The main types:

1. Primary processing

  • Data cleaning (removing duplicates, correcting errors)
  • Data filtering
  • Normalisation and standardisation
  • Merging data includes the integration, consolidation and cleaning of data before it is analysed.
  • Data validation checks and ensures the correctness of input data before it is used further. Validation is performed before data is saved or processed. The main goal — is to make sure that the data meets the requirements (for example, that an email has the correct format, that a number does not exceed the permissible limits). If data fails validation, it is discarded or corrected, which is a key part of data cleaning.

  • Data hydration — is the process of converting raw data from a database into entity objects. The reverse process (converting objects into an array or an SQL query) is called dehydration (Dehydration). Hydration = turning SQL data into PHP objects

2. Analytical processing

  • Statistical analysis
  • Machine learning and neural networks
  • Forecasting and modelling
  • Data distillation (Data Distillation) – is the process of extracting the most significant information from large volumes of data. It is closer to analytical processing, since it is used in machine learning and for creating simplified data models.

3. Logical processing

  • Finding patterns
  • Classification and categorisation
  • Sorting – “arranging items in a particular sequence and/or into various sets”.
  • Optimisation and decision-making
  • ORM translation = turning DQL queries into SQL queries. Converting object-oriented queries (DQL) into SQL. Before the query is executed, when Doctrine converts DQL into SQL.

4. Data visualisation

  • Building charts and diagrams
  • Interactive dashboards
  • GIS maps (geographic information systems)

5. Storage, transfer, destruction and management of data

  • Structuring (SQL, NoSQL)
  • Archiving and backups
  • Ensuring data security
  • transfer is the process of moving information from one source to another through various communication channels (the internet, a local network, Bluetooth, radio waves, etc.).

  • Data destruction — is the process of deleting information in such a way that it cannot be recovered. This is important for ensuring data security, complying with information-protection laws, and also preventing leaks of confidential data.

Information and data processing

Means of information processing

With the current state of software development there exist many different software tools for processing information, written in various programming languages, based on the methods listed above. The diversity of software is linked to the specifics of each field in which processing is carried out. For example, when processing graphic images, pattern-recognition methods and cryptographic methods based on the Fourier transform, among others, are widely used.

Among the information-processing tools available to a broad range of users are tools for organising databases, executing queries and searching for information, filtering information, graphical presentation, and so on.

At the current stage, methods of human-oriented computer data processing are increasingly being developed.

At present, owing to the global spread of computer systems in the field of industrial process automation, systems for data acquisition and operational dispatch control (SCADA — Supervisory Control And Data Acquisition System) are being used ever more widely. SCADA is only one component of automated control systems, which at the current stage form a complex system of software and hardware. The overwhelming majority of automated control systems are built on the basis of industrial controllers, which serve as the primary means of collecting and processing information, regulating process parameters, alarm signalling, protection and interlocking (the lower level of the system). Information processed by the controllers is transmitted to a computerised system that serves as the operator-technologist's workstation, where further processing of process data takes place and it is presented to the operator in an intuitively understandable form (the upper level of the automated process control system). In the hierarchy of software and hardware for industrial automation, SCADA systems occupy the upper level. A SCADA system collects information about the process, provides an interface with the operator, stores the history of the process, and exercises control over the process to the extent necessary.

Automated information processing

The operational capabilities of the modern complex of technical means used in the system of automated collection and processing of information make it possible to automatically carry out a whole range of procedures within these functions. The state of scientific and practical developments and the technical level of the aforementioned complex have determined the possibilities for the automated performance of the following management-process procedures:

  • in forecasting and planning — multivariate calculations in developing forecasts, long-term and current economic and social development plans for the enterprise, as well as operational-production plans and plans for the technical preparation of production, with a view to further determining optimal interrelated sets of planning indicators in hourly (hour, shift, week, etc.) and per-object (workplace, section, etc.) terms;
  • in organisation — modelling organisational management structures and simulating production processes under different criteria and parameters with a view to choosing the optimal ones;
  • in coordination and regulation — issuing commands to workplaces (at the lower level of production management) in accordance with the plan, the technological process or instructions drawn up for particular kinds of work or operations;
  • in control — monitoring the state of the managed object across all parameters, as well as the timely and complete execution of management commands;
  • in accounting — the one-time collection (in step with production) and systematic processing of all actual (together with reference, planning, normative and other) reliable information on the availability and movement of resources, as well as on the states, processes and phenomena taking place in the production, economic and other activity of the enterprise;
  • in analysis — comparing normative, planned and actual indicators characterising particular operations or processes of production, economic or other activity, identifying deviations (in quantitative, monetary, relative and other terms) from the specified parameters with an indication of the causes and those responsible for these deviations, evaluating the fulfilment of the plan from various angles and identifying the factors affecting these deviations;
  • in reporting — the automatic generation (on the basis of primary data, of summary indicators for standard forms of established accounting, statistical and other reporting, using special conversion arrays) of reference books, — as well as the simultaneous creation of machine media with summary reporting indicators for transmission over communication channels to external institutions (agencies) at a higher level.

History

The history of the United States Census Bureau illustrates the evolution of data processing from manual to electronic procedures.

Manual data processing

Although the widespread use of the term “data processing” dates only from the 1950s, data-processing functions were performed manually for thousands of years. For example, accounting includes such functions as posting transactions and compiling reports, such as a balance sheet and a cash-flow statement . Fully manual methods were supplemented by the use of mechanical or electronic calculators . A person whose job consisted of performing calculations by hand or with the aid of a calculator was called a “ computer ”.

The 1890 United States census schedule was the first in which data was collected on individual persons rather than by household . A number of questions could be answered by ticking the appropriate box on the form. From 1850 to 1880 the Census Bureau used “a tallying system which, owing to the increasing number of classification combinations required, was becoming ever more complex. Only a limited number of combinations could be recorded in a single tally, so the schedules had to be processed 5 or 6 times for the same number of independent tallies” “It took more than 7 years to publish the results of the 1880 census” using manual-processing methods.

Automatic data processing

The term “automatic data processing” was applied to operations performed using unit-record equipment, for example to Herman Hollerith's application of punched-card equipment to the 1890 United States census . “Using Hollerith's punched-card equipment, the Census Bureau was able to complete the tabulation of most of the 1890 census data in 2–3 years, compared with 7–8 years for the 1880 census. It is estimated that the use of the Hollerith system saved about 5 million dollars in processing costs” in 1890 dollars, even though there were twice as many questions as in 1880.

Computerised data processing

Computerised data processing, or electronic data processing, is a later development in which a computer was used instead of several independent pieces of equipment. The Census Bureau first used a limited number of electronic computers for the 1950 United States census, using the UNIVAC I system, delivered in 1952.

Other developments

The term data processing was largely subsumed into the more general term information technology (IT). The older term “data processing” implies older technology. For example, in 1996 the Data Processing Management Association (DPMA) changed its name to the Association of Information Technology Professionals . Nevertheless, these terms are roughly synonymous.

See also

  • Big data
  • Computing
  • Computer science
  • Decision support software
  • Information age
  • Information and communication technology
  • Information technology
  • Scientific computing
  • Data analysis
  • Databases
  • Data warehouses
  • Knowledge bases
  • Algorithm

created: 2025-01-31
updated: 2026-03-10
109



Was this answer useful?
Choose a quick rating so we can improve the next answer for you.
How satisfied are you?


Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Data mining"

Terms: Data mining