South 고등학교 2학년 Data Science
Chapters: 4
1. Understanding data science
What data science is · Structured and unstructured data · Databases · Data and social change
- Data Science: From a Question to a Tested Model – Data science is using data, maths and computers to answer questions and make predictions. A project follows a cycle: ask a clear question, collect data, clean it (fix or remove errors), explore it with charts and averages, build a model, test the model, and report the result. A model is a simple rule learned from data, such as marks ≈ 35 + 7 × hours. Prediction (regression) models give a number; classification models give a group. We test models on new data and measure accuracy or error, then choose the best one. Data science is used in health, sport, weather, shops and farming. Real projects often fail at first, so patience and checking matter.
- Data Handling: Collect, Organise and Show Data – Data is a set of facts, such as answers, counts or measurements. Data can be qualitative (words) or quantitative (numbers); numbers are discrete (counted) or continuous (measured). We collect data by surveys, observation, experiments or from existing sources, then organise it with tally marks into a frequency table. We show it with the right graph: bar graphs to compare groups, pie charts to show parts of a whole, line graphs for change over time and scatter graphs for links between two variables. Computers store data as structured tables or unstructured text, images and sound.
- Database Concepts: DBMS, Relations and Keys – Keeping data in separate files causes duplication, inconsistency and poor security. A database stores related data in one organised place, and a DBMS (like MySQL) is the software that manages it. In the relational model, data is kept in tables (relations) made of columns (attributes) and rows (tuples); a domain is the set of allowed values of a column. A candidate key uniquely identifies each row; one is chosen as the primary key and the others are alternate keys.
- Digital Society: How Information Technology Changes Our Lives – A digital society (also called an information society) is a society where making, sharing and using information with computers and networks is a main part of life, work and government. Information moves fast, reaches everywhere and can be copied at almost no cost. Software has changed shopping, banking, health, school and friendship. New jobs appear (app developer, data analyst) while some old jobs shrink. Every online action leaves data, so privacy matters. Not everyone is connected (the digital divide), and too much screen time or gaming can harm health. A good digital citizen uses the benefits and manages the risks.
2. Preparing and analysing data
Unbiased data collection · Outliers, missing values, normalisation · Transforming data for analysis · Different analyses of the same data
- Sampling: Learning About a Population from a Sample – A population is the whole group we want to know about; a sample is a smaller part we actually check. A good sample is chosen at random so that it represents the population. Simple random sampling gives everyone an equal chance; stratified sampling takes the right share from each group; systematic sampling takes every k-th item. Different samples give slightly different answers (sampling variation), but bigger samples wobble less (law of large numbers). A biased sample gives a wrong answer however big it is.
- Data Analysis – Data analysis means turning raw data into answers. It follows a cycle: ask a question, collect data, clean it (remove errors, repeats and blanks), organise and transform it, analyse it with summaries such as mean, median, range and patterns, show it with a good chart, and draw a careful conclusion. Watch for outliers, small samples and bias, and remember that a correlation between two things does not prove that one causes the other. Data must also be stored safely and used with permission.
3. Modelling and evaluation
Data models · Regression models · Similarity and clustering · Correlation and relationships · Model outputs by method · Comparing and evaluating models
- Data Science: From a Question to a Tested Model – Data science is using data, maths and computers to answer questions and make predictions. A project follows a cycle: ask a clear question, collect data, clean it (fix or remove errors), explore it with charts and averages, build a model, test the model, and report the result. A model is a simple rule learned from data, such as marks ≈ 35 + 7 × hours. Prediction (regression) models give a number; classification models give a group. We test models on new data and measure accuracy or error, then choose the best one. Data science is used in health, sport, weather, shops and farming. Real projects often fail at first, so patience and checking matter.
- Linear Regression and the Least Squares Line – Linear regression finds the straight line ŷ = a + bx that best follows paired data (x, y). A residual is the gap between a real point and the line: e = y − ŷ. The least squares line makes the sum of squared residuals as small as possible. Its slope is b = Sxy ÷ Sxx and it always passes through the mean point (x̄, ȳ). We use it to predict y from x, but only inside the data range, and a strong link does not prove that x causes y.
- Correlation: Scatter Diagram, Karl Pearson's Coefficient and Spearman's Rank Correlation – Correlation tells how two variables move together. It is positive when both rise together, negative when one rises as the other falls, and zero when there is no straight-line pattern. A scatter diagram shows it as a picture. Karl Pearson's coefficient r measures its direction and strength and always lies between −1 and +1. Spearman's rank correlation R uses ranks and works for qualities like beauty or honesty; tied ranks need a small correction.
4. Data science project
Applications by field · Exploratory data analysis · Combining several datasets · Persistence with hard problems · Full project cycle
- Data Science: From a Question to a Tested Model – Data science is using data, maths and computers to answer questions and make predictions. A project follows a cycle: ask a clear question, collect data, clean it (fix or remove errors), explore it with charts and averages, build a model, test the model, and report the result. A model is a simple rule learned from data, such as marks ≈ 35 + 7 × hours. Prediction (regression) models give a number; classification models give a group. We test models on new data and measure accuracy or error, then choose the best one. Data science is used in health, sport, weather, shops and farming. Real projects often fail at first, so patience and checking matter.
- Data Analysis – Data analysis means turning raw data into answers. It follows a cycle: ask a question, collect data, clean it (remove errors, repeats and blanks), organise and transform it, analyse it with summaries such as mean, median, range and patterns, show it with a good chart, and draw a careful conclusion. Watch for outliers, small samples and bias, and remember that a correlation between two things does not prove that one causes the other. Data must also be stored safely and used with permission.