South 고등학교 3학년 Mathematics for Artificial Intelligence
Chapters: 4
1. AI and mathematics
Maths behind AI history · How AI uses maths
- Artificial Intelligence: How Machines Learn to Think – Artificial intelligence (AI) is the skill of a computer system to do tasks that normally need human thinking: seeing, understanding speech, deciding and learning. An AI system is an agent that senses, thinks and acts. Old AI followed rules written by people. Modern AI mostly uses machine learning: it finds its own rule from many labelled examples (data). Neural networks are layers of simple units whose link strengths (weights) change during training. AI is used in maps, translation, health, farming and games. It can be wrong or unfair when its data is one-sided (bias), so people must check it, protect privacy and stay responsible.
2. Representing data
Text as numbers · Processing text data · Images as numbers · Processing image data
- Vector Algebra – A vector has a size (magnitude) and a direction. In 3D we write it as a = xî + yĵ + zk̂. Its length is |a| = √(x² + y² + z²). Its direction cosines are l = x/|a|, m = y/|a|, n = z/|a|, and l² + m² + n² = 1. Vectors are added head-to-tail (triangle law) or component by component. ka stretches a by k and flips it if k is negative. The point dividing AB in m : n has position vector (mb + na)/(m + n) inside and (mb − na)/(m − n) outside. Dot product a·b = |a||b|cosθ gives a number and tells the angle and the projection. Cross product a×b = |a||b|sinθ n̂ gives a vector at right angles to both; its length is the area of the parallelogram on a and b.
- Matrices: Order, Types, Transpose, Operations and Inverse – A matrix is a box of numbers set in rows and columns. Its order is rows × columns. Special matrices include zero, identity, diagonal, scalar, row, column and square matrices. The transpose swaps rows and columns. Symmetric means Aᵀ = A and skew-symmetric means Aᵀ = −A. We add matrices place by place, and multiply them row × column. Matrix multiplication is not commutative (AB is usually not BA). A square matrix A is invertible if some B gives AB = BA = I, and that B is unique.
3. Classification and prediction
Classifying text · Classifying images · Predicting with probability · Trend lines
- Machine Learning – Machine learning (ML) is a way for computers to learn a rule from examples instead of being told the rule. We give data made of features (inputs) and, in supervised learning, labels (answers). The machine fits a model: regression predicts a number, a decision tree asks yes/no questions to pick a class, and k-means clustering groups unlabelled data. We train on most of the data, test on data it never saw, and measure accuracy or error. A model that only memorises (overfits) fails on new data.
- Conditional Probability, Multiplication Rule and Independent Events – Conditional probability is the chance of A when we already know B has happened. We throw away every outcome outside B and count again: P(A|B) = P(A ∩ B) ÷ P(B). Turned around, this gives the multiplication rule P(A ∩ B) = P(B)·P(A|B). If knowing B does not change the chance of A, the events are independent and P(A ∩ B) = P(A)·P(B).
- Linear Regression and the Least Squares Line – Linear regression finds the straight line ŷ = a + bx that best follows paired data (x, y). A residual is the gap between a real point and the line: e = y − ŷ. The least squares line makes the sum of squared residuals as small as possible. Its slope is b = Sxy ÷ Sxx and it always passes through the mean point (x̄, ȳ). We use it to predict y from x, but only inside the data range, and a strong link does not prove that x causes y.
4. Optimisation
Loss functions · Finding minima · Rational decisions
- Linear Regression and the Least Squares Line – Linear regression finds the straight line ŷ = a + bx that best follows paired data (x, y). A residual is the gap between a real point and the line: e = y − ŷ. The least squares line makes the sum of squared residuals as small as possible. Its slope is b = Sxy ÷ Sxx and it always passes through the mean point (x̄, ȳ). We use it to predict y from x, but only inside the data range, and a strong link does not prove that x causes y.
- Gradient Descent: Finding the Lowest Point – Gradient descent finds the minimum of a function by taking small steps downhill. At each step, new x = old x − learning rate × slope. A learning rate that is too big overshoots; too small is very slow. It is how machine learning models reduce their error.
- Decision Making: How to Choose Well, Step by Step – Decision making means choosing the best option from two or more choices. A good decision follows steps: spot the problem, list options, set criteria, give each criterion a weight, score the options, choose, act and then review. A decision matrix turns this into simple numbers.