Lecture
The sample (empirical) distribution function in mathematical statistics is an approximation of the theoretical distribution function, constructed using a sample drawn from it.
Let be a sample of size
, generated by a random variable
, given by the distribution function
. We will assume that
, where
, are independent random variables defined on some sample space of elementary outcomes
. Let
. Define the function
as follows:
,
where is the indicator of event
,
is the Heaviside function. Thus, the value of the function
at the point
equals the relative frequency of sample elements not exceeding the value
. The function
is called the sample distribution function of the random variable
, or the empirical function of the sample, and is an approximation of the function
. There is a Kolmogorov theorem, which asserts that as
the function
converges uniformly to
, and which indicates the rate of convergence. For every positive
,
is a random variable with value
.
Since the unknown distribution
can be described, for example, by its distribution function
, let us construct, from the sample, an «estimate» for this function.
The empirical distribution function, constructed from the sample
of size
, is the random function
, which for each
equals


is called the indicator of the event
. For each
this is a random variable having a Bernoulli distribution with parameter
.
In other words, for any
the value
, equal to the true probability that the random variable
is less than
, is estimated by the fraction of sample elements less than
.
If the sample elements
,
,
are ordered in increasing order (for each elementary outcome), a new set of random variables is obtained, called the order statistics (variational series):

Here

The element
,
, is called the
-th term of the order statistics or the
-th order statistic.
,
where , and
is the number of sample elements equal to
. In particular, if all sample elements are distinct, then
.
The expected value of this distribution has the form:
.
Thus, the sample mean is the theoretical mean of the sample distribution. Similarly, the sample variance is the theoretical variance of the sample distribution.
.
.
.
almost surely as
.
in distribution as
.
Sample: 
Order statistics: 

The empirical distribution function has jumps at the sample points; the size of the jump at the point
equals
, where
is the number of sample elements coinciding with
.
The empirical distribution function can be constructed from the order statistics as follows:

Another characteristic of a distribution is the table (for discrete distributions) or the density (for absolutely continuous ones). The empirical, or sample, analog of the table or the density is the so-called histogram.
A histogram is constructed from grouped data. The presumed range of values of the random variable
(or the range of the sample data) is divided, independently of the sample, into a certain number of intervals (not necessarily equal). Let
,
,
be intervals on the line, called the grouping intervals. For
denote by
the number of sample elements falling into the interval
:
(1)
Over each of the intervals
a rectangle is constructed, whose area is proportional to
. The total area of all rectangles must equal one. Let
be the length of the interval
. The height
of the rectangle over
equals

The resulting figure is called a histogram.
We have the order statistics (see Example 1):

Let us divide the segment
into 4 equal segments. Into the segment
fell 4 sample elements, into
— 6, into
— 3, and into the segment
fell 2 sample elements. We construct the histogram (Fig. 2). In Fig. 3 there is also a histogram for the same sample, but with the range divided into 5 equal segments.

Fig. 2, Fig. 3
In the course «Econometrics» it is asserted that the best number of grouping intervals («Sturges' formula») is
.
Here
is the decimal logarithm, so
, i.e., when the sample size is doubled, the number of grouping intervals increases by 1. Note that the more grouping intervals there are, the better. But if one takes the number of intervals, say, of order
, then as
grows the histogram will not approach the density.
The following statement holds:
If the density of the distribution of the sample elements is a continuous function, then as
in such a way that
, pointwise convergence in probability of the histogram to the density takes place.
So the choice of the logarithm is reasonable, but it is not the only possible one.
Comments