📘 CodingMarble Learn

Data Science: From a Question to a Tested Model

Data science is using data, maths and computers to answer questions and make predictions. A project follows a cycle: ask a clear question, collect data, clean it (fix or remove errors), explore it with charts and averages, build a model, test the model, and report the result. A model is a simple rule learned from data, such as marks ≈ 35 + 7 × hours. Prediction (regression) models give a number; classification models give a group. We test models on new data and measure accuracy or error, then choose the best one. Data science is used in health, sport, weather, shops and farming. Real projects often fail at first, so patience and checking matter.

🎬 Step-by-step story

  1. Every data science project goes round a cycle. Ask a question. Collect data. Clean it. Explore it. Build a model. Report. Then ask again.
  2. Each blue dot is one student: hours studied and marks. The red dots say 140 and minus 10 marks. Those are impossible, so we clean them away.
  3. Now a line runs through the middle of the dots. This line is the model. It predicts that 6 hours of study gives about 77 marks.
  4. A different model sorts students into groups instead of giving a number. Green means pass, orange means at risk. The black line is the border.
  5. We test the model on 20 new students. 17 boxes are green (right) and 3 are red (wrong). Accuracy is 17 out of 20, which is 85%.
  6. Your turn. Drag the slope slider. Watch the average error. Find the line that sits closest to the dots.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Is data science just making charts?

No. Charts are one step (explore). Data science goes all the way from a question to a tested model and a report. See the full ring in step 1.

Why not keep every data point?

Impossible values come from mistakes and pull the model the wrong way. Step 2 shows the red points being removed.

How does a line 'predict'?

You read up from the hours to the line and across to the marks. The yellow dot in step 3 shows 6 hours → about 77 marks.

What is the difference between regression and classification?

Regression gives a number (step 3); classification gives a group (step 4, green or orange).

How do we know if a model is good?

Test it on new data and count how many it gets right. Step 5 shows 17 out of 20.

What is the 'best fit' line?

The line with the smallest average distance to the points. Find it with the slider in step 6.

What is data science?

Data are facts we record: numbers, words, pictures, clicks. Data science is the skill of turning data into useful answers using statistics (maths of data), computing and knowledge of the topic.

Data has value when it helps someone decide better: what to stock, when to water crops, which patient to see first.

A short history of managing data

People first kept data on clay tablets and paper ledgers. Then came punched cards (around 1890), computer databases (1960s–70s), spreadsheets (1980s), the internet and big data (2000s), and now machine learning, where computers learn patterns from huge datasets.

The data science project cycle

  1. Ask: write a clear question. "Does study time affect marks?"
  2. Collect: survey, sensors, records, open datasets.
  3. Clean: remove impossible values, fix typing errors, handle missing data, remove duplicates.
  4. Explore: draw charts, find averages and patterns.
  5. Model: build a rule that explains or predicts.
  6. Report: share findings with charts and simple words, and say what the limits are.

The cycle repeats: the answer often leads to a better question.

Models and what they output

A model is a simplified rule learned from data.

To read a model's output, ask: what does each number mean, and in what units? A slope of 7 means each extra hour adds about 7 marks.

Comparing and evaluating models

Always test a model on new data it did not learn from.

The better model has higher accuracy or lower error on new data. A model that is perfect on old data but poor on new data has overfitted: it memorised instead of learning.

Uses, ethics and persistence

Uses: health (spotting disease), sport (player tactics), farming (rain and pest alerts), transport (traffic), business (stock and prices), science (climate studies).

Ethics: protect privacy, ask permission, and check the data is fair. A model trained on biased data gives biased answers.

Persistence: real data is messy. Models often fail at first. Good data scientists try again, test ideas one at a time and write down what they learn.

Try it at home

Ask 10 friends how many hours they slept last night and how alert they feel (score 1–10). Write the data in a table. Remove anything impossible (like 30 hours). Plot the points on squared paper and draw a line through the middle. Use your line to predict the score for 7 hours of sleep. Then ask two more friends and check how close you were.

Key formulas and definitions

Worked examples

1. Use the model marks ≈ 35 + 7 × hours to predict marks for 4 hours of study.

35 + 7 × 4 = 35 + 28 = 63 marks.

2. A spam filter checks 200 emails and gets 184 right. What is its accuracy?

184 ÷ 200 × 100% = 92%.

3. A model predicts 60, 72 and 80 marks. The real marks were 64, 70 and 85. Find the average error.

Errors: |64 − 60| = 4, |70 − 72| = 2, |85 − 80| = 5. Average = (4 + 2 + 5) ÷ 3 = 11 ÷ 3 ≈ 3.7 marks.

4. Model A has 90% accuracy on training data and 70% on new data. Model B has 82% and 80%. Which should we use?

Model B. It works almost as well on new data. Model A has overfitted: it memorised the old data.

Common mistakes

Practice quiz

1. Which comes first in a data science project?
2. Removing impossible values and duplicates is called…
3. A model that says 'spam' or 'not spam' is a…
4. 17 right out of 20 gives an accuracy of…
5. Overfitting means a model…

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is data science in simple words?

Using data, maths and computers to answer questions and make predictions that help people decide.

What are the steps of the data science process?

Ask a question, collect data, clean it, explore it, build and test a model, then report the findings.

What is the difference between data science and big data?

Big data means very large, fast and varied datasets. Data science is the method of getting answers from data of any size.

Where this is taught

South Korea고등학교 2학년Understanding data science
South Korea고등학교 2학년Modelling and evaluation
South Korea고등학교 2학년Data science project
China高二Sel.3 Data management and analysis

Learn first

Learn next

Related lessons

All Computer Science lessons