📘 CodingMarble Learn

Inferential Statistics: Significance, Errors and Choosing a Test

Inferential statistics tell us whether a result in a sample is likely to be real or just chance. We start from a null hypothesis (no effect) and ask how likely our data would be if it were true. If that probability is 5% or less (p ≤ 0.05), we call the result significant and reject the null hypothesis. We can still be wrong: a Type I error is a false positive, a Type II error is a false negative. The right test depends on three things: whether we look for a difference or a correlation, whether the design is related or unrelated, and the level of data (nominal, ordinal or interval). We then compare the calculated value with a critical value from a table.

đŸŽŦ Step-by-step story

  1. The null hypothesis says "no real effect, just chance". Bars show how likely each result is when only chance is at work.
  2. If a result is so rare that chance explains it 5% of the time or less (p ≤ 0.05), we call it significant.
  3. We can make two mistakes. Type I: we say there is an effect when there is none. Type II: we miss a real effect.
  4. To choose a test, ask two questions: what is the design, and what is the level of data? One box gives the test.
  5. Worked sign test: count the plus and minus signs, drop the zeros, take the less common sign as S, and compare S with the critical value.
  6. Free play: choose a design and a level of data. The box with the right test rises up.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why start from "no effect" instead of from our hypothesis?

We can work out how likely data are under pure chance, so we test against that and see if our result is too rare. Step 1 shows the chance-only bars.

Does significant mean the effect is big?

No. It means the result is unlikely to be chance. A tiny effect can be significant with a large sample. Step 2 marks only the rare zone.

Can we ever be sure there is no error?

No. There is always some risk of a Type I or Type II error. Step 3 shows all four outcomes.

How do I know if data are ordinal or interval?

If the units are equal and standard (seconds, cm), they are interval. If they are ratings or ranks, they are ordinal. Try both rows in the free play.

Why do we drop the zeros in a sign test?

A zero shows no change in either direction, so it gives no evidence for or against. Step 5 drops the grey ball.

What do the 9 boxes of the grid mean?

Rows are the level of data, columns the design. Where they cross is your test. Step 4 shows the grid.

Probability and significance

Probability (p) is how likely something is, from 0 (never) to 1 (certain). In psychology we use it to judge whether a result could be just chance.

Type I and Type II errors

The 5% level is a balance between the two.

Choosing a test

Ask three questions:

  1. Difference or correlation?
  2. Design: unrelated (independent groups) or related (repeated measures or matched pairs)?
  3. Level of data: nominal (categories, counts), ordinal (can be ranked, gaps not equal, e.g. ratings) or interval (equal units on a public scale, e.g. time in seconds).
UnrelatedRelatedCorrelation
NominalChi-squaredSign testChi-squared
OrdinalMann-Whitney (U)Wilcoxon (T)Spearman's rho
IntervalUnrelated t-testRelated t-testPearson's r

Parametric tests (the t-tests and Pearson) also need data that are roughly normally distributed with similar spread in both groups.

Calculated and critical values: a worked sign test

Each test gives a calculated (observed) value. You compare it with a critical value from a table. To find the right critical value you need: the significance level, one- or two-tailed, and N (number of participants) or df (degrees of freedom).

Example: 12 people rate their mood before and after a walk. 9 feel better (+), 2 feel worse (−), 1 is the same (0).

  1. Drop the zero: N = 11.
  2. S = the number of the less common sign = 2.
  3. The hypothesis was directional, so use one-tailed p = 0.05. For N = 11 the critical value is 2.
  4. Sign test has no R, so S must be ≤ critical value. 2 ≤ 2, so the result is significant. We reject the null hypothesis.

Try it: a coin test at home

Flip a coin 10 times and count heads. Do this 5 times. How often did you get 0, 1, 9 or 10 heads? Those results are in the red "rare" zone of step 2. Then pick a study idea, decide its design and level of data, and check your test choice in the 3D free play.

Key formulas and definitions

Worked examples

1. Researchers compare reaction times (in milliseconds) of two separate groups. Which test?

Difference, unrelated design, interval data → unrelated t-test.

2. Do students' ranks in maths and in music relate? Data are ranks. Which test?

Correlation, ordinal data → Spearman's rho.

3. A chi-squared test gives a calculated value of 5.2; the critical value is 3.84. Is it significant?

Chi-squared has an R, so the calculated value must be â‰Ĩ critical. 5.2 â‰Ĩ 3.84, so yes, significant at p ≤ 0.05.

Common mistakes

Practice quiz

1. The usual significance level in psychology is:
2. Rejecting a true null hypothesis is a:
3. Ranked data where gaps between ranks are not equal are:
4. Related design, ordinal data, looking for a difference:
5. In the sign test, S is:

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is the difference between Type I and Type II errors?

Type I is a false positive: rejecting a true null hypothesis. Type II is a false negative: keeping a false null hypothesis.

How do you choose a statistical test in psychology?

Decide if you want a difference or a correlation, whether the design is related or unrelated, and whether the data are nominal, ordinal or interval. The grid then gives the test.

What is a critical value?

The value from a table that the calculated value is compared with to decide whether a result is significant, based on N or df, the significance level and one- or two-tailed.

Where this is taught

England (GCSE, A level)Year 134.2.3 Research methods (A-level extension)

Learn first

Learn next

Related lessons

All Psychology lessons