Distribution

Distribution

Distribution Definition

In quantitative research, a distribution is the pattern formed by all the observed values of a variable — whether measures of individual income, city crime levels, or health differences between nations. The statistician’s approach to examining a distribution is the same regardless of what is measured or whether the units are individuals or countries. As a frequency (or statistical) distribution, it is a series of figures presenting every observed value, as raw numbers or proportions of cases, enabling a quick visual appreciation of the data, often illustrated with a pie chart or histogram.

Four Aspects to Examine

When examining any distribution there are four key aspects to consider; missing any of them risks overlooking important features of the data.

Central tendency summarizes the average value. Most commonly this is the mean — the sum of all cases divided by their number — though in many situations the median, the value of the middle case when all are rank-ordered, gives a better typical value. For categorical variables the mode, the most frequently occurring category, is the usual measure.

Spread describes how heterogeneous the cases are, and sociologists are often more interested in it than in averages. Among the richest few dozen countries, for instance, quality of life varies surprisingly little with mean income, but the spread of incomes — the gap between richest and poorest — appears to have a greater effect on outcomes such as health and average life expectancy. Researchers frequently pay too much attention to averages and neglect spread. Common measures include the standard deviation and the midspread (interquartile range).

Shape matters because many statistical tests assume data plotted as a histogram will form a bell-shaped curve — the normal or Gaussian distribution. In practice few sociological variables do, making “normal” something of a misnomer. Hourly income in most countries is highly skewed, with a long upward straggle toward a small number of very high earners while most wages sit slightly below the mean. Weekly working hours form a “bimodal” graph, with one peak around full-time (36–40 hours) and another around half-time (20 hours). In such cases the shape tells far more than the average.

Outliers are the small number of cases that differ sharply from all the rest. Counting sexual partners over twelve months, most people would score 0 or 1, but a few — prostitutes, for example — would report dozens or hundreds; pooling all cases into an average would mislead. For some analyses it is appropriate to exclude such extremes; in others, the deviant cases are the most instructive — the exceptions that prove the rule. But beware: extreme cases often arise from errors in the research itself.

Observed versus Probability Distributions

An observed frequency distribution should not be confused with mathematical probability distributions, which are hypothetical, their form determined by algebraic formulae. The correspondence between an observed distribution and various hypothetical ones often determines which statistical analysis is appropriate. Frequency distributions from a survey dataset are usually the first output produced from the cleaned and edited data, showing response totals for every possible answer to each questionnaire item; they can then be analyzed using measures of dispersion and other tools developed from the three main probability distributions — binomial, Poisson, and normal (Gaussian).

Sociology Plus
Logo