Probability and significance
Probability (p) is how likely something is, from 0 (never) to 1 (certain). In psychology we use it to judge whether a result could be just chance.
- The null hypothesis says there is no difference or no relationship.
- The alternative (experimental) hypothesis says there is one. It is directional (one-tailed) if it predicts which way, and non-directional (two-tailed) if it does not.
- The usual significance level is p ⤠0.05: there is 5% or less probability the result happened by chance if the null hypothesis is true.
- A stricter level, p ⤠0.01, is used when a mistake would be costly, for example in medical research or when replicating a surprising result.
Type I and Type II errors
- Type I error (false positive): we reject a null hypothesis that is actually true. We claim an effect that is not there. More likely if the significance level is too lenient (for example p ⤠0.10).
- Type II error (false negative): we keep a null hypothesis that is actually false. We miss a real effect. More likely if the level is too strict (for example p ⤠0.01) or the sample is small.
The 5% level is a balance between the two.
Choosing a test
Ask three questions:
- Difference or correlation?
- Design: unrelated (independent groups) or related (repeated measures or matched pairs)?
- Level of data: nominal (categories, counts), ordinal (can be ranked, gaps not equal, e.g. ratings) or interval (equal units on a public scale, e.g. time in seconds).
| Unrelated | Related | Correlation | |
|---|---|---|---|
| Nominal | Chi-squared | Sign test | Chi-squared |
| Ordinal | Mann-Whitney (U) | Wilcoxon (T) | Spearman's rho |
| Interval | Unrelated t-test | Related t-test | Pearson's r |
Parametric tests (the t-tests and Pearson) also need data that are roughly normally distributed with similar spread in both groups.
Calculated and critical values: a worked sign test
Each test gives a calculated (observed) value. You compare it with a critical value from a table. To find the right critical value you need: the significance level, one- or two-tailed, and N (number of participants) or df (degrees of freedom).
Example: 12 people rate their mood before and after a walk. 9 feel better (+), 2 feel worse (â), 1 is the same (0).
- Drop the zero: N = 11.
- S = the number of the less common sign = 2.
- The hypothesis was directional, so use one-tailed p = 0.05. For N = 11 the critical value is 2.
- Sign test has no R, so S must be ⤠critical value. 2 ⤠2, so the result is significant. We reject the null hypothesis.
Try it: a coin test at home
Flip a coin 10 times and count heads. Do this 5 times. How often did you get 0, 1, 9 or 10 heads? Those results are in the red "rare" zone of step 2. Then pick a study idea, decide its design and level of data, and check your test choice in the 3D free play.
Key formulas and definitions
- Significant if p ⤠0.05 (5% or less probability of chance)
- Type I = false positive; Type II = false negative
- Sign test: S = count of the less common sign; N excludes zeros
- R rule: R in the name â calculated âĨ critical; no R â calculated ⤠critical
Worked examples
1. Researchers compare reaction times (in milliseconds) of two separate groups. Which test?
Difference, unrelated design, interval data â unrelated t-test.
2. Do students' ranks in maths and in music relate? Data are ranks. Which test?
Correlation, ordinal data â Spearman's rho.
3. A chi-squared test gives a calculated value of 5.2; the critical value is 3.84. Is it significant?
Chi-squared has an R, so the calculated value must be âĨ critical. 5.2 âĨ 3.84, so yes, significant at p ⤠0.05.
Common mistakes
- Saying "significant" means "important". It only means unlikely to be chance.
- Saying p ⤠0.05 proves the hypothesis. It only lets us reject the null hypothesis, with a 5% risk of a Type I error.
- Counting the zeros in a sign test. Zeros are dropped and N gets smaller.
- Using the R rule backwards: for sign test, Wilcoxon and Mann-Whitney the calculated value must be less than or equal to the critical value.