What are descriptive statistics, and how do I report them in my thesis?

Before you test hypotheses, run regressions or build structural equation models, you need to describe your data. Who are your respondents? What are the typical values of your key variables? How spread out are the scores? Are there unusual values or missing data? Descriptive statistics answer these questions.

Descriptive statistics may seem basic, but they are essential. They give readers a clear picture of your sample and variables, help you check data quality and assumptions, and provide context for interpreting inferential results. Examiners often look closely at the descriptive section to judge whether you understand your data. This post explains the main descriptive statistics, when to use each one, and how to report them clearly in your thesis.

What are descriptive statistics?

Descriptive statistics summarise and describe the main features of a dataset. They do not, on their own, test hypotheses or generalise beyond the sample. That is the role of inferential statistics, such as t-tests, correlations and regression.

Descriptive statistics usually cover three areas:

  • Frequency distributions: how often each value or category occurs.
  • Measures of central tendency: the typical or central value.
  • Measures of variability (dispersion): how spread out the values are.

They may also include measures of distribution shape, such as skewness and kurtosis.

Frequency distributions

Frequencies show how many cases fall into each category. They are most useful for categorical variables such as gender, education level, department or employment status.

Report both the count (n) and the percentage (%). For example: “Of the 250 respondents, 142 (56.8%) were women, 104 (41.6%) were men, and 4 (1.6%) preferred not to say.”

For continuous variables, frequencies can be displayed as grouped categories (such as age bands) or with a histogram.

Measures of central tendency

Mean

The mean is the arithmetic average: the sum of all values divided by the number of values. It is appropriate for interval and ratio data that are roughly symmetrically distributed. The mean is sensitive to extreme values (outliers), which can pull it up or down.

Median

The median is the middle value when data are ordered from lowest to highest. It is less affected by outliers and skewed distributions, so it is often preferred for variables such as income, reaction time or length of hospital stay, which tend to be skewed.

Mode

The mode is the most frequent value. It is most useful for nominal data, such as the most common job category.

Which one should I use?

  • Nominal data: mode and frequencies.
  • Ordinal data: median (and sometimes mode); for single Likert items, frequencies and median are often recommended.
  • Interval/ratio data, roughly symmetric: mean.
  • Interval/ratio data, skewed or with outliers: median, often with the mean as well.

Scale scores created by averaging several Likert items are commonly reported with means and standard deviations in social science research.

Measures of variability

Range

The difference between the highest and lowest values. It is simple but heavily influenced by extreme values. Reporting the minimum and maximum is often more informative than the range alone.

Standard deviation (SD)

The standard deviation shows how much values typically differ from the mean. A small SD means values cluster close to the mean; a large SD means they are more spread out. The SD is reported alongside the mean.

Variance

The variance is the square of the standard deviation. It is important in statistical calculations but is less intuitive to report, so SD is usually preferred in descriptive tables.

Interquartile range (IQR)

The IQR is the range of the middle 50% of values (from the 25th to the 75th percentile). It is reported with the median for skewed data.

Distribution shape

Skewness

Skewness describes asymmetry. Positive skew means a long tail to the right (a few high values); negative skew means a long tail to the left.

Kurtosis

Kurtosis describes how heavy the tails of the distribution are compared with a normal distribution.

Skewness and kurtosis values are often used to assess approximate normality before parametric tests. Different authors propose different thresholds for acceptable values, so cite the guideline you use. Visual checks with histograms and Q–Q plots are also recommended, because formal normality tests such as Shapiro–Wilk can be very sensitive in large samples.

Describing your sample

Most theses include a table describing participant characteristics, sometimes called a demographic profile. Include variables relevant to your research questions and to judging representativeness, such as:

  • Age (mean, SD and range, or age groups).
  • Gender.
  • Education level.
  • Occupation, role or years of experience.
  • Organisation type or location.

Where possible, compare your sample with known population characteristics to discuss representativeness.

Describing your main study variables

For your key variables, such as scale scores, provide a table with:

  • Number of valid cases (n).
  • Mean and standard deviation (or median and IQR if appropriate).
  • Minimum and maximum values, or the possible range of the scale.
  • Skewness and kurtosis, if relevant to your analysis.
  • Reliability (such as Cronbach's alpha), often included in the same table.

Many theses also include a correlation matrix in the same or an adjacent table, showing relationships between all main variables.

Example of a descriptive statistics table

Below is an example layout for a table of main study variables. The values are illustrative only.

Table 4.2 Descriptive Statistics and Reliability for Main Study Variables (N = 312)

VariableItemsRangeMSDSkewnessKurtosisα
Job satisfaction51–53.620.78−0.410.12.88
Work–life conflict61–52.610.840.910.65.85
Supervisor support41–74.951.21−0.58−0.20.91
Turnover intention31–52.341.020.49−0.53.83

Note. M = mean; SD = standard deviation; α = Cronbach's alpha. Higher scores indicate higher levels of each construct.

A table like this lets readers see at a glance how each variable was measured, what the typical scores were, how much variation there was, and whether the scales were reliable.

Running descriptive statistics in SPSS

In SPSS, the most commonly used menus are:

  • Analyze → Descriptive Statistics → Frequencies: for categorical variables, and for detailed statistics (including median, quartiles, skewness and kurtosis) and histograms for continuous variables.
  • Analyze → Descriptive Statistics → Descriptives: for a quick table of N, mean, SD, minimum and maximum for several continuous variables.
  • Analyze → Descriptive Statistics → Explore: for detailed statistics by group, boxplots, normality plots and tests.
  • Analyze → Descriptive Statistics → Crosstabs: for frequencies across two categorical variables.

Always paste and save your syntax so that you can reproduce the output later. Then build your thesis tables from the output rather than copying screenshots.

Example of reporting in text

Here is an example in a style common in social sciences (APA style), with illustrative numbers:

“Participants (N = 312) were aged between 22 and 61 years (M = 38.4, SD = 9.7). Most were women (n = 187, 59.9%) and held a bachelor's degree or higher (n = 254, 81.4%). Job satisfaction scores ranged from 1.4 to 5.0 on a 5-point scale (M = 3.62, SD = 0.78). Work–life conflict scores were moderately positively skewed (skewness = 0.91), so medians are also reported (Mdn = 2.40, IQR = 1.80–3.20).”

Notice that the text highlights key points rather than repeating every number from the table.

Formatting tips

  • Follow your required style guide, such as APA, Harvard or your university's thesis guidelines.
  • Use consistent decimal places, usually one or two for means and SDs.
  • Italicise statistical symbols such as M, SD, n and N in APA style.
  • Use tables for detail and text for highlights. Do not repeat all table values in the text.
  • Label tables clearly, with a number and descriptive title, and define abbreviations in a note below the table.
  • Report valid n for each variable if there are missing data.
  • Include the scale range so readers can interpret means, such as “1 = strongly disagree to 5 = strongly agree”.

Using descriptive statistics to check data quality

Descriptive statistics are also a diagnostic tool. Before analysis, use them to:

  • Check for impossible or out-of-range values (for example, a score of 7 on a 1–5 scale).
  • Identify missing data patterns.
  • Spot outliers that may affect results.
  • Check distribution shapes for assumption testing.
  • Confirm that reverse-coded items have been recoded correctly.

Report how you handled any problems, such as removing invalid cases or treating missing data.

Common mistakes to avoid

  • Reporting means for nominal variables, such as “mean gender = 1.4”.
  • Reporting only means without standard deviations.
  • Using means for highly skewed data without comment.
  • Inconsistent decimal places or missing sample sizes.
  • Pasting raw software output into the thesis instead of formatted tables.
  • Repeating every table value in the text.
  • Interpreting descriptive differences as if they were statistically tested.

Final thoughts

Descriptive statistics are the foundation of your results chapter. They show readers who your participants are, what your key variables look like, and whether your data are ready for further analysis. Choose the right measures for each variable type, present them in clear, well-formatted tables, highlight the key points in the text, and use descriptive analysis to check data quality. A strong descriptive section builds confidence in everything that follows.

Have questions about which descriptive statistics to report for your variables? Ask in the comments below.