Data Definition
Data (the plural of datum, Latin for “given”) are facts or observations — the pieces of information from which inferences are drawn through analysis. They take many forms: responses to questionnaires and interviews, IQ scores, fieldwork diaries, taped interviews, observational and documentary records, statistics, visual materials, and oral histories. Although qualitative researchers regard their observations as no less factual than survey statistics, the term often connotes quantifiable information — hence “databases” and “data banks.” Strictly, data is a plural noun (“the data show”), though singular usage is increasingly common. The Latin sense of “things given” is itself misleading, since gathering evidence usually demands painstaking effort; the researcher’s key question is how these sources of knowledge relate to theory, the source of understanding.
Quantitative and Qualitative Data
Data design — deciding how best to collect evidence capable of answering the research question — often begins with the distinction between quantitative and qualitative data. The two camps have frequently been set against each other as though they held incompatible views of how to know the social world, but the overlap is considerable, and the real aim is a good fit between question and data type. Pay differences between men and women in the United Kingdom cannot be established without statistics, yet qualitative material on how men and women view their pay adds insight; explaining women’s under-representation in senior management requires both statistical evidence on job applicants’ success rates and richly contextualised information about men’s and women’s differing life circumstances. Ideally the two complement each other.
Primary and Secondary Data
Social research is not solely first-hand collection: vast quantities of information already exist, gathered by others and often for other purposes. Official statistics, collected by government agencies, vary in how “official” they are — vital statistics are by-products of the statutory registration of births, marriages, and deaths, while many national surveys rest on voluntary participation. The Census is the most extensive source, and longitudinal studies such as Britain’s ONS Longitudinal Study (based on one per cent of the population) track change year by year. Growing computing power has revolutionised archiving and dissemination, allowing researchers to download whole surveys or sub-samples for secondary analysis — the extraction of new findings by reanalysing existing data. Users of secondary data must, however, stay mindful of the original purpose and agendas behind the collection: statistics are to some extent social constructions whose concepts reflect dominant official, political, and economic categories. Employment surveys, for instance, may define work in ways that ignore unpaid care and so undervalue women’s economic contribution — a limitation to weigh, not a reason to discard the data.
Documentary Sources
Written documents likewise offer great opportunities but can rarely be taken at face value; assessing them requires knowing why they were produced. Documents differ by authorship (personal or official) and by access (from closed to openly published). John Scott’s A Matter of Record (1990) argues that such knowledge underpins four key questions: a document’s authenticity (is it original and genuine?), credibility (is it accurate?), representativeness (does it stand for the totality of documents of its class?), and meaning (what is it intended to say?).
Ethics and Data Abuse
Data cannot be separated from ethical and political concerns. Professional bodies’ guidelines govern collection — stipulating practices such as informed consent — but arguably the gravest abuses occur at the analysis and writing-up stages, when claims outrun what the evidence warrants. W. G. Runciman, in The Social Animal (1998), calls this an “abuse of social science,” citing sociologists who espouse ideology or opinion far beyond what their data support; the warning applies whatever the political complexion of the conclusions. Data archives, both quantitative and qualitative, help guard against such failings: they widen access for secondary research, foster the cumulative character of social science by enabling replication and extension of earlier work, and safeguard professional standards by keeping evidence available for scrutiny and reanalysis.

