Correlation analysis is one of the most common statistical procedures in thesis research. It answers a simple question: are two variables related, and if so, how strongly and in which direction? For example, is job satisfaction related to turnover intention? Is study time related to exam performance? Is social media use related to loneliness?
SPSS makes running a correlation easy, but interpreting and reporting it correctly requires understanding what the output means, which type of correlation to choose, and what assumptions to check. This post walks you through the whole process step by step, from preparing your data to writing up your results.
What does a correlation tell you?
A correlation coefficient summarises the strength and direction of the relationship between two variables.
- Direction: a positive coefficient means both variables tend to increase together; a negative coefficient means one tends to decrease as the other increases.
- Strength: the coefficient ranges from −1 to +1. Values closer to −1 or +1 indicate stronger relationships; values near 0 indicate weak or no linear relationship.
- Statistical significance: the p-value tells you whether the observed correlation is unlikely to have occurred by chance if there were no relationship in the population.
Remember that correlation shows association, not causation.
Choosing the right type of correlation
Pearson's correlation (r)
Pearson's r measures the strength of a linear relationship between two continuous variables (interval or ratio). It is appropriate when:
- Both variables are continuous (or treated as approximately continuous, such as scale scores averaged from several Likert items).
- The relationship is approximately linear.
- There are no extreme outliers strongly influencing the result.
- For significance testing, the variables are approximately normally distributed, particularly in smaller samples.
Spearman's rank-order correlation (rho, rs)
Spearman's rho is based on ranks rather than raw values. It measures the strength of a monotonic relationship (one that consistently increases or decreases, but not necessarily in a straight line). Use it when:
- One or both variables are ordinal, such as single Likert items or rankings.
- Data are clearly non-normal or have outliers.
- The relationship is monotonic but not linear.
Kendall's tau-b
Another rank-based correlation, sometimes preferred for small samples or data with many tied ranks.
Step 1: Prepare your data
Before running correlations:
- Check that variables are correctly coded and missing values are defined in SPSS (Variable View → Missing).
- Reverse-code negatively worded items and compute scale scores if needed (Transform → Compute Variable).
- Check reliability of your scales.
- Run descriptive statistics to check ranges and spot errors.
Step 2: Check assumptions
Linearity and outliers
Create a scatter plot for each key pair of variables: Graphs → Chart Builder → Scatter/Dot (or Graphs → Legacy Dialogs → Scatter/Dot → Simple Scatter). Look for:
- A roughly linear pattern (for Pearson).
- Curved patterns, which suggest Pearson may underestimate the relationship.
- Outliers that sit far from the main cluster of points.
Normality
Use Analyze → Descriptive Statistics → Explore to obtain histograms, Q–Q plots, skewness and kurtosis for each variable. If variables are strongly skewed, consider Spearman's rho, or report both Pearson and Spearman as a robustness check.
Step 3: Run the correlation in SPSS
- Go to Analyze → Correlate → Bivariate.
- Move the variables you want to correlate into the Variables box. You can include several variables to produce a correlation matrix.
- Under Correlation Coefficients, tick Pearson, Spearman or Kendall's tau-b as appropriate.
- Under Test of Significance, choose Two-tailed unless you have a strong, pre-specified directional hypothesis and your field accepts one-tailed tests.
- Tick Flag significant correlations if you want SPSS to mark them with asterisks.
- Click Options to request means and standard deviations, and choose how to handle missing values: Exclude cases pairwise (uses all available data for each pair) or Exclude cases listwise (uses only cases with complete data on all variables).
- Recent versions of SPSS also offer options for confidence intervals for correlations; if available, request them, as they are increasingly expected in reporting.
- Click Paste to save the syntax, then run it.
Step 4: Read the output
The output shows a Correlations table. For each pair of variables, you will see:
- Pearson Correlation (or Correlation Coefficient for Spearman): the coefficient r (or rs).
- Sig. (2-tailed): the p-value.
- N: the number of cases used for that pair.
The table is symmetrical: the correlation between A and B appears twice. The diagonal shows each variable correlated with itself (always 1).
Note that when SPSS shows “.000” for Sig., it means p is less than .001, not exactly zero. Report it as “p < .001”.
Step 5: Interpret the results
Direction
A positive r means higher values on one variable are associated with higher values on the other. A negative r means higher values on one are associated with lower values on the other.
Strength
Cohen (1988) offered widely cited rough guidelines for r in the behavioural sciences:
- About .10: small.
- About .30: medium.
- About .50: large.
Cohen himself warned that these are general conventions, not rigid rules. What counts as a meaningful relationship depends on your field and context. Compare your results with previous research on similar variables.
Shared variance
Squaring r gives the coefficient of determination (r²), the proportion of variance in one variable that is statistically associated with the other. For example, r = .40 means r² = .16, or 16% shared variance.
Significance
A significant p-value (for example, p < .05) suggests the correlation is unlikely to be zero in the population. But significance depends heavily on sample size. In large samples, very small correlations can be significant; in small samples, moderate correlations may not reach significance. Always interpret the size of the coefficient, not only the p-value.
Step 6: Report the results
In text (APA style example)
“A Pearson correlation analysis showed that job satisfaction was negatively correlated with turnover intention, r(248) = −.46, p < .001, 95% CI [−.55, −.36], indicating a moderate-to-large relationship. Employees with higher job satisfaction tended to report lower intentions to leave.” (Numbers are illustrative.)
Note that the number in brackets after r is the degrees of freedom, which equals N − 2 for a Pearson correlation.
For Spearman: “Spearman's rank-order correlation indicated a positive relationship between frequency of mentoring meetings and perceived career support, rs(98) = .31, p = .002.”
In a correlation matrix table
When reporting many correlations, present a matrix with variable names in rows and numbered columns, means and SDs in the first columns, and correlations in the lower triangle. Mark significance with asterisks and explain them in a table note, such as “*p < .05. **p < .01.”. Reliability coefficients are sometimes shown on the diagonal.
Example correlation matrix
Here is an example of how a formatted correlation table might look in a thesis. Values are illustrative only.
Table 4.3 Means, Standard Deviations and Pearson Correlations Among Study Variables (N = 250)
| Variable | M | SD | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|---|
| 1. Job satisfaction | 3.58 | 0.81 | (.89) | |||
| 2. Supervisor support | 3.72 | 0.92 | .52** | (.91) | ||
| 3. Workload | 3.41 | 0.77 | −.29** | −.18** | (.84) | |
| 4. Turnover intention | 2.47 | 1.05 | −.46** | −.33** | .24** | (.87) |
Note. Cronbach's alpha values are shown in parentheses on the diagonal. **p < .01 (two-tailed).
Interpreting a matrix like this
When you describe a correlation matrix in your results chapter, do not describe every single coefficient. Instead, highlight the relationships that matter for your research questions and hypotheses. For the example above, you might write:
“As expected, job satisfaction was strongly and positively correlated with supervisor support (r = .52) and negatively correlated with turnover intention (r = −.46). Workload showed small-to-moderate correlations with the other variables, being negatively related to job satisfaction (r = −.29) and positively related to turnover intention (r = .24). All correlations were statistically significant at p < .01. None of the correlations between predictor variables exceeded .60, suggesting that multicollinearity is unlikely to be a serious concern for the subsequent regression analysis.”
This kind of paragraph links the correlations to your hypotheses, comments on strength rather than just significance, and prepares the reader for the next stage of analysis. The comment about multicollinearity should be supported by formal checks, such as variance inflation factors, when you run the regression itself.
Partial correlation
Sometimes you want to know whether two variables are related after controlling for a third. SPSS offers Analyze → Correlate → Partial. For example, you might examine the correlation between training hours and performance while controlling for years of experience. Partial correlation can be a useful step, although multiple regression is often used when several control variables are involved.
Common mistakes to avoid
- Using Pearson's r without checking scatter plots for non-linearity or outliers.
- Reporting only p-values without the correlation coefficient.
- Reporting “p = .000”.
- Interpreting significant correlations as causal relationships.
- Running many correlations without considering the risk of chance findings.
- Correlating single Likert items with Pearson's r without justification.
- Pasting raw SPSS output into the thesis instead of a formatted table.
Final thoughts
Running a correlation in SPSS takes only a few clicks, but meaningful analysis requires more: choosing the correct coefficient, checking assumptions with scatter plots and descriptive statistics, interpreting both the size and significance of the relationship, and reporting results clearly and honestly. Follow these steps and you will produce correlation results that are accurate, well-presented and easy for examiners to trust.
Do you have a question about your SPSS correlation output? Share it in the comments below.