Dummy Variables

Dummy Variables

Dummy Variables Definition

Dummy variables are devices used in quantitative data analysis to allow variables that are not measured at the interval level to be included in regression analysis. Researchers are often interested in the implications of non-interval variables for a dependent variable — the relationship between sex and income, for example — yet regression normally requires interval-level data. Creating dummy variables solves the problem.

How They Are Constructed

In the simplest case of a two-category variable such as sex, one category is coded 1 and the other 0 — men as 1 and women as 0, for instance. Where an independent variable comprises more than two categories — say n categories — it is represented by n − 1 dummy variables. If “social class” contains the four categories upper, middle, working, and none, three dummies are created: the first coded 1 for “upper” and 0 for “not upper”; the second coded 1 for “middle” and 0 for “not middle”; the third coded 1 for “working” and 0 for “not working.” The fourth category needs no dummy of its own, since it is fully described by the combination 0-0-0. Every category thus has a unique combination of zeros and ones by which its presence or absence is indicated.

Use in Regression

Once the dummies are constructed, the categorical information can enter a multiple regression alongside genuinely interval-level predictors, and the resulting regression coefficients are treated as if they were based on interval-level variables. Each coefficient expresses the effect on the dependent variable of membership in that category relative to the omitted reference category — in the example above, the effect of being upper, middle, or working class relative to having no class assignment. The technique is a standard part of survey analysis and connects to broader issues of measurement in quantitative sociology.

Sociology Plus
Logo