United Grade 12 AP Statistics
Chapters: 5
1. Exploring One-Variable Data and Collecting Data
Introducing Statistics: What Can We Learn from Data? · Variables · Tabular Representation and Summary Statistics for One Categorical Variable · Graphical Representations for One Categorical Variable · Graphical Representations for One Quantitative Variable · Descriptions for One Quantitative Variable Distributions · Summary Statistics for One Quantitative Variable · Graphical Representations of Summary Statistics for One Quantitative Variable · Comparisons of the Distributions for One Quantitative Variable · The Investigative Question Revisited and Data Collection · Random Sampling · Potential Problems with Sampling · Experimental Design
- Measures of Dispersion: Range, Mean Deviation, Variance and SD – Dispersion means spread: how far the values sit from the centre. Range = largest − smallest. Mean deviation = average distance from the mean (or median). Variance = average of squared distances from the mean. Standard deviation = √variance. The same ideas work for grouped data when every term is multiplied by its frequency.
- Sampling: Learning About a Population from a Sample – A population is the whole group we want to know about; a sample is a smaller part we actually check. A good sample is chosen at random so that it represents the population. Simple random sampling gives everyone an equal chance; stratified sampling takes the right share from each group; systematic sampling takes every k-th item. Different samples give slightly different answers (sampling variation), but bigger samples wobble less (law of large numbers). A biased sample gives a wrong answer however big it is.
- The Scientific Method – The scientific method is the careful way scientists find out how the world works. Observe something, ask a testable question, make a hypothesis (a clear, testable guess), test it with a fair experiment (change one variable, measure one, keep the rest the same), repeat and record data, analyse it, draw a conclusion and share it so others can check. Results that fail the test are useful too: they send you back to a new hypothesis.
2. Probability, Random Variables, and Probability Distributions
Tabular and Graphical Representations for the Distributions of Two Categorical Variables · Summary Statistics for Two Categorical Variables · Estimating Probabilities Using Simulation · Introduction to Probability · Mutually Exclusive Events · Conditional Probability · Independent Events and Unions of Events · Introduction to Random Variables and Probability Distributions · Parameters of Random Variables · The Binomial Distribution · The Normal Distribution · Sampling Distributions and the Central Limit Theorem
- Two-Way Tables – A two-way table counts data for two categorical variables at once: one in the rows, one in the columns. Row and column totals are marginal totals. Dividing a cell by the grand total gives a joint relative frequency; dividing by its row (or column) total gives a conditional relative frequency. If conditional percentages differ a lot between groups, the variables are associated.
- Estimating Probability Using Simulation – A simulation imitates a chance process with random numbers so we can estimate a probability that is hard to calculate. Plan it: state the question, give each outcome a set of random digits that matches its probability, define one trial and what counts as the event, then repeat many trials. The relative frequency (event count ÷ trials) is the estimate, and it gets closer to the true probability as the number of trials grows (the law of large numbers).
- Probability: Events, Algebra of Events and Axioms – An event is any subset of the sample space S. From events A and B we build new events: not A (A′), A and B (A ∩ B), A or B (A ∪ B). Events are mutually exclusive if they share no outcome and exhaustive if together they cover S. The axiomatic approach says every P(E) ≥ 0, P(S) = 1, and for mutually exclusive A, B, P(A ∪ B) = P(A) + P(B). From these follow P(A′) = 1 − P(A) and P(A ∪ B) = P(A) + P(B) − P(A ∩ B).
- Conditional Probability, Multiplication Rule and Independent Events – Conditional probability is the chance of A when we already know B has happened. We throw away every outcome outside B and count again: P(A|B) = P(A ∩ B) ÷ P(B). Turned around, this gives the multiplication rule P(A ∩ B) = P(B)·P(A|B). If knowing B does not change the chance of A, the events are independent and P(A ∩ B) = P(A)·P(B).
- Random Variables and Probability Distributions – A random variable X gives a number to every outcome of a chance experiment. A discrete random variable takes separate values; its probability distribution lists each value x with its probability P(X = x), and these add up to 1. The mean (expected value) is E(X) = Σ x·P(x). The variance Var(X) = E(X²) − [E(X)]² measures spread, and the standard deviation is its square root.
- Binomial Distribution – Repeat the same yes/no trial n times, independently, with the same chance of success p each time. The number of successes X follows the binomial distribution B(n, p): P(X = k) = ⁿCₖ pᵏ qⁿ⁻ᵏ with q = 1 − p. Its mean is np and its variance is npq.
- Normal Distribution – A normal distribution is a continuous, symmetric, bell-shaped distribution fixed by its mean μ (centre) and standard deviation σ (spread): X ~ N(μ, σ²). Mean = median = mode. About 68% of values lie within 1σ of μ, 95% within 2σ and 99.7% within 3σ. Any normal value is turned into a standard score z = (x − μ)/σ, which follows N(0, 1); probabilities are areas under the curve, read from a table or calculator as Φ(z).
- Sampling Distributions and the Central Limit Theorem – A statistic (like a sample mean x̄ or a sample proportion p̂) changes from sample to sample. If you took every possible sample and plotted the statistic, you would get its sampling distribution. Its centre is the true population value (μ or p), its spread is the standard error (σ/√n for means, √(p(1−p)/n) for proportions), and by the Central Limit Theorem its shape becomes close to normal when n is large enough.
3. Inference for Categorical Data: Proportions
Estimators · Sampling Distributions for Sample Proportions · Constructing a Confidence Interval for a Population Proportion · Justifying a Claim Based on a Confidence Interval for a Population Proportion · Setting Up a Test for a Population Proportion · p-Values · Carrying Out a Test for a Population Proportion · Potential Errors When Performing Tests · Sampling Distributions for the Difference Between Sample Proportions · Constructing a Confidence Interval for the Difference Between Two Population Proportions · Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions · Setting Up a Test for the Difference Between Two Population Proportions · Carrying Out a Test for the Difference Between Two Population Proportions · Setting Up a Chi-Square Test for Homogeneity or Independence · Carrying Out a Chi-Square Test for Homogeneity or Independence
- Sampling Distributions and the Central Limit Theorem – A statistic (like a sample mean x̄ or a sample proportion p̂) changes from sample to sample. If you took every possible sample and plotted the statistic, you would get its sampling distribution. Its centre is the true population value (μ or p), its spread is the standard error (σ/√n for means, √(p(1−p)/n) for proportions), and by the Central Limit Theorem its shape becomes close to normal when n is large enough.
- Confidence Intervals – A confidence interval is a range of believable values for an unknown population number (a mean μ or a proportion p), worked out from one sample. It has the shape estimate ± margin of error, where margin of error = critical value × standard error. A 95% level means the method catches the true value in about 95% of samples. Higher confidence gives a wider interval; a bigger sample gives a narrower one. If a claimed value lies outside the interval, the data give evidence against the claim.
- Hypothesis Testing – A hypothesis test checks a claim about a population using a sample. Start with the null hypothesis H₀ (no change, e.g. p = 0.5) and the alternative H₁ (what we suspect, e.g. p > 0.5). Choose a significance level such as 5%. Work out how likely the sample result (or more extreme) is if H₀ were true: the p-value. If the p-value is below the level, or the result falls in the critical region, reject H₀. Otherwise there is not enough evidence to reject it. Type I error = rejecting a true H₀; Type II error = not rejecting a false H₀.
- Chi-Square Test for Independence and Homogeneity – A chi-square (χ²) test checks if counts in a table are too far from what we would expect by chance. For a contingency table: E = row total × column total ÷ grand total, χ² = Σ (O − E)² ÷ E, df = (r − 1)(c − 1). If χ² is bigger than the critical value (or p < significance level), reject H0 of no association. For 2×2 tables, Yates' correction uses (|O − E| − 0.5)².
4. Inference for Quantitative Data: Means
Sampling Distributions for Sample Means · Constructing a Confidence Interval for a Population Mean or Population Mean Difference · Justifying a Claim Based on a Confidence Interval for a Population Mean or Population Mean Difference · Setting Up a Test for a Population Mean or Population Mean Difference · Carrying Out a Test for a Population Mean or Population Mean Difference · Sampling Distributions for the Difference Between Two Sample Means · Constructing a Confidence Interval for the Difference Between Two Population Means · Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means · Setting Up a Test for the Difference Between Two Population Means · Carrying Out a Test for the Difference Between Two Population Means
- Sampling Distributions and the Central Limit Theorem – A statistic (like a sample mean x̄ or a sample proportion p̂) changes from sample to sample. If you took every possible sample and plotted the statistic, you would get its sampling distribution. Its centre is the true population value (μ or p), its spread is the standard error (σ/√n for means, √(p(1−p)/n) for proportions), and by the Central Limit Theorem its shape becomes close to normal when n is large enough.
- Confidence Intervals – A confidence interval is a range of believable values for an unknown population number (a mean μ or a proportion p), worked out from one sample. It has the shape estimate ± margin of error, where margin of error = critical value × standard error. A 95% level means the method catches the true value in about 95% of samples. Higher confidence gives a wider interval; a bigger sample gives a narrower one. If a claimed value lies outside the interval, the data give evidence against the claim.
- Hypothesis Testing – A hypothesis test checks a claim about a population using a sample. Start with the null hypothesis H₀ (no change, e.g. p = 0.5) and the alternative H₁ (what we suspect, e.g. p > 0.5). Choose a significance level such as 5%. Work out how likely the sample result (or more extreme) is if H₀ were true: the p-value. If the p-value is below the level, or the result falls in the critical region, reject H₀. Otherwise there is not enough evidence to reject it. Type I error = rejecting a true H₀; Type II error = not rejecting a false H₀.
5. Regression Analysis
Graphical Representations Between Two Quantitative Variables · Correlation · Linear Regression Models · Residuals · Least-Squares Regression
- Correlation: Scatter Diagram, Karl Pearson's Coefficient and Spearman's Rank Correlation – Correlation tells how two variables move together. It is positive when both rise together, negative when one rises as the other falls, and zero when there is no straight-line pattern. A scatter diagram shows it as a picture. Karl Pearson's coefficient r measures its direction and strength and always lies between −1 and +1. Spearman's rank correlation R uses ranks and works for qualities like beauty or honesty; tied ranks need a small correction.