Lecture
A causal relationship is a connection between phenomena in which one phenomenon, called the cause, under certain conditions gives rise to another phenomenon, called the effect.
If a dog was given meat at the same time a light bulb was switched on, then after several repetitions the dog began salivating not only at the sight of the meat itself, but also at the switching on of the light bulb. A conditioned reflex was formed. The repetition of the coincidence of two stimuli is the cause, the reflex is the effect.
A common mistake is to see a causal relationship in a correlation.
The greatest life expectancy is observed in the regions of Scotland with the lowest population density and the lowest unemployment rate. In the USA, life expectancy correlates with income level (the lives of the poor and people of low socioeconomic status more often end prematurely). In modern Great Britain, professional status correlates with life expectancy. According to the results of a study conducted over 10 years involving 17,350 British civil servants, the number of deaths among managerial staff is 1.6 times higher, and among clerical workers and laborers — 2.2 and 2.7 times higher, respectively, than among senior management (Adler et al., 1993, 1994). It appears that at different times and in different geographic locations, there is a fairly definite interdependence between status and health.
The example above of the relationship between status and life expectancy illustrates one of the most common thinking errors, among both amateurs and professionals: when two factors «go hand in hand», such as status and health, it is hard to resist the temptation to conclude that one is the cause of the other. One might assume that status somehow protects a person from things that could harm their health. Or is it not like that at all, and good health is not an effect but a cause of activity and success? Perhaps long-livers manage to accumulate more money, and that is precisely why their graves have more expensive headstones? Correlational research allows us to make a prediction, but it cannot answer the question of whether a change in one parameter (for example, social status) will cause a change in another parameter (for example, health status).
Fact: children who were frequently beaten by their parents usually do worse in school and more often display antisocial behavior. Does this mean that one follows from the other? It is not at all obvious. Rather, both the harsh punishment of children and poor academic performance together with antisocial behavior are a consequence of the fact that these children grew up in troubled families.
Fact: children with a developed sense of self-worth usually do better in school than children with low self-esteem. Does this mean that self-worth is the cause and good academic performance is the effect? No, correlation here says nothing yet about which is the cause and which is the effect
Causal research , also called explanatory research , - is research into cause-and-effect relationships ( a study ) . To determine a causal relationship, it is important to observe a change in the variable that is presumed to cause a change in another variable (or variables), and then measure the changes in the other variable (or variables). Other confounding influences must be controlled so that they do not distort the results, either by holding them constant when experimentally generating data, or by using statistical methods. This type of research is very complex, and the researcher can never be fully certain that there are no other factors influencing the causal relationship, especially when it comes to people's relationships and motivation. There are often much deeper psychological considerations that even the respondent may not be aware of.
There are two research methods for studying the causal relationship between variables: experimentation (for example, in a laboratory ) and statistical research .
Experiments are usually conducted in laboratories, where many or all aspects of the experiment can be strictly controlled in order to avoid false results due to factors other than the intended causal factor(s). For example, many studies in physics use this approach. Alternatively, field experiments can be conducted, as in medical research, in which subjects may have very many attributes that cannot be controlled, but in which, at least, key hypothetical causal variables can be varied, and some of the extraneous attributes can at least be measured. . Field experiments are also sometimes used in economics , for example, when two different groups of social benefit recipients are given two alternative sets of incentives or income-earning opportunities, and their effect on labor supply is investigated.
In fields such as economics , most empirical research is conducted on pre-existing data that is often collected by the government on a regular basis. Multiple regression is a group of related statistical methods that control for (attempt to avoid the false influence of) various correlations other than the one being studied. If the data show sufficient variation in the hypothetical explanatory variable of interest, its correlation with the potentially affected variable can be measured. This, however, does not imply causation.
How did you first discover that a light bulb turns on if you flip the switch? How do you know that a gun, when fired, produces a loud sound, and not the other way around?
We gain knowledge of causes in two main ways:
Perception (causal experience). Seeing a brick fly through a window, one billiard ball strike another, causing it to roll, a lit match ignite a candle wick, we form impressions of causal dependence based on incoming sensory information.
Inference (mediated conclusions about causation using the deductive method and based on non-causal information). The causes of events such as food poisoning, wars, and good health cannot be perceived directly — they must be inferred through logical reasoning based on something other than direct observation.
The trust we place in causal perception can let us down. If you hear a loud sound, and afterward the light turns on in the room, it is easy to decide that these events are related; however, the temporal coincidence of the loud sound and the moment someone flips the switch may be a simple coincidence.
Trust in causal perception can let us down.
The temporal and spatial proximity of events — are parameters because of which we often draw false conclusions.
For example, we often hear that a person got a flu vaccination, and by evening they developed flu-like symptoms, and people believe that it was the shot that caused this. But the flu vaccine, containing an inactive form of the virus, cannot cause the illness. Among the huge number of vaccinated people, some develop other similar illnesses (by pure coincidence), or they catch the virus while waiting to be seen at the clinic.
Time
Events close together in time can lead to erroneous conclusions about causation. Imagine: you have a headache and you took some remedy. A few hours later, the pain went away. Can it be claimed that the medicine helped?
The temporal pattern allows us to assume that the symptom eased thanks to taking the medicine, but you cannot say for certain that the pain would not have gone away on its own. You would have to conduct many randomized experiments, where you would take or not take the drug, and then record how quickly the headache disappeared, in order to be able to claim anything at all about such a causal dependence. You would also have to compare the effects of the drug and a placebo.
Causal dependence cannot always be justified.
Long delays between cause and effect can also hinder the reliable establishment of causal relationships. Some effects occur quickly (a strike on a billiard ball makes it move), while some processes proceed in slow motion. It is known that smoking causes lung cancer; but many long years lie between the first cigarette and the day the cancer is diagnosed.
Side effects from taking certain drugs appear decades later. Changes in health due to physical exercise are achieved slowly and not immediately, and if we focus only on the scale's reading, it may seem that weight even increases at first, because muscle is built up faster than fat is lost. Expecting the effect to follow immediately after the cause, we fail to see the connection between these deeply interdependent factors.
Correlation
Correlation (a relationship, an association) does not necessarily mean causal dependence. This idea is firmly drilled into the mind of any statistics student; but sometimes even those who understand and agree with this statement make mistakes.
A strong relationship may seem convincing and prompt a series of successful predictions. But apparent correlations are sometimes explained by causes that have not yet been measured.
For example, we found a relationship in a situation where a person who ate a hearty breakfast makes it to work on time; however, most likely, both factors have a common cause: the person got up early, and therefore had time to have a good breakfast, instead of rushing to work.
Correlation does not necessarily mean causal dependence.
Having identified a correlation between two variables, one needs to check whether such an unmeasured factor (a common cause) can explain this relationship.
Moreover, relationships can exist even when two variables are not connected in any way at all. Correlations can result from pure chance (for example, you run into a friend on the street many times in a week), artificial experimental conditions (questions may be tailored to specific responses), or an error or malfunction (a bug in a computer program).
Without variation there is no correlation
Imagine this situation: you want to find out how to get a grant, so you ask all your friends who have received one what, in their opinion, helped them. All the candidates prepared their application in Times New Roman font; according to half of them, it is important that each page have at least one illustration; and a third recommend submitting the application 24 hours before the deadline. Does this mean there is a correlation between the stated conditions and receiving the grant? No, it does not.
Since all the outcomes are identical, one cannot say what would happen if the font were changed or the application were submitted a minute before the deadline.
Without variation there is no correlation.
And yet it is a widespread situation where only the factors leading to a particular outcome are analyzed. Just imagine how often winners are asked exactly how they achieved success, and then people try to reproduce that success by performing precisely the same actions.
Such an approach is full of flaws for many reasons, including that people simply are not very good at identifying the relevant factors, underestimate the role of chance, and overestimate their own abilities. As a result, we not only confuse factors that purely by chance accompany the desired effect with those that actually produce it, but we also see illusory correlations where none exist.
People are not very good at identifying the relevant factors, underestimate the role of chance, and overestimate their own abilities. Source
Talking to winners is useless, since one could do the exact same thing and not succeed. Perhaps all candidates prepare their grant applications in Times New Roman font (meaning those who did not receive grants would recommend using a different font), or perhaps successful candidates received the grant despite having an excessive number of illustrations in their documents. Without knowing the full set of positive and negative examples, we cannot even assume the presence of a correlation.
Selection bias
One important reason we make erroneous conclusions is that the data may not be representative of the underlying distribution.
If we were only allowed to look at flu death statistics, but were given only the data on the number of patients admitted to medical facilities, we would observe a much higher percentage of fatal outcomes than across the entire population. This happens because people end up in hospital, as a rule, with more severe cases or additional illnesses (and with high chances of dying from the flu). Thus we are comparing not all outcomes, but only the statistics for those who sought medical help with flu-like symptoms.
Selection data must be representative.
Or take, for example, websites that survey visitors about their political views. On the internet, it is not possible to randomly select survey participants representative of the entire population, and data from sources with a strong political bias are distorted even more.
If the visitors to a particular page actively support the incumbent president, the results among them might show that the head of state's approval rating rises every time he gives an important speech. However, this only shows that there is a correlation between the president's approval and his giving speeches to supporters.
Confirmation bias
Some of the cognitive biases that make us see a relationship between unrelated factors are similar to selection bias. For example, confirmation bias makes us look for evidence in favor of a particular belief.
In other words, if you believe that a drug causes a certain side effect, you will start reading online reviews from those who have already taken it and observed that effect. But in doing so, you ignore the entire set of data that does not support your hypothesis, instead of looking for evidence that might make you reassess it.
Confirmation bias can also make you dismiss evidence that contradicts your hypothesis; you might assume that the source of the information is unreliable or that the study was based on flawed experimental methods.
Confirmation bias.
In addition to bias in terms of evidence, an error in interpreting arguments can occur. If, during "non-blind" testing of a new drug, a doctor remembers that a patient is taking this remedy and believes it is helping them, they may start looking for signs of its effectiveness. Since many parameters are subjective (for example, mobility or fatigue), this can lead to distortions in the assessment of these indicators and to logical conclusions about the existence of non-existent correlations.
There is also a specific form of confirmation bias — illusory correlation. It refers to looking for a relationship where none exists. The possible connection between arthritis symptoms and weather is so widely publicized that it is considered proven. However, knowing about it can lead patients to report a correlation simply because they expect to see it. When scientists tried to analyze this issue, using patient reports, clinical tests, and objective indicators as a basis, they found absolutely no connection whatsoever.
Mathematics
Physics
Philosophy
Philosophy of mind
Statistics
Psychology and medicine
Pathology and epidemiology
Sociology and economics
Environmental issues
Comments