📘 CodingMarble Learn

Data Science Methodology: From a Problem to a Working Model

Data science methodology is a step-by-step cycle for solving a problem with data. Its 10 stages are: business understanding (the problem), analytic approach, data requirements, data collection, data understanding, data preparation, modelling, evaluation, deployment and feedback. Data is split into training and test sets (or checked with k-fold cross-validation) so the model is tested on data it has not seen. Classification models are judged by accuracy; regression models by errors such as MSE and RMSE. Feedback from real use starts the cycle again.

🎬 Step-by-step story

  1. Every data project starts with a clear problem. Then we choose the type of answer we need: a class, a number or groups.
  2. Next we decide what data we need, collect it and look at it carefully to understand it.
  3. Real data is messy. In preparation we remove duplicates, fill or drop missing values and fix wrong entries.
  4. In modelling we split the data. The model learns from the training part. The test part stays hidden.
  5. We evaluate the model on the test data, deploy it for real use, and use feedback to improve it. Then the cycle repeats.
  6. Free play: pick any stage to see what happens there, and change the train-test split.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Why not start with the model straight away?

Without a clear question you do not know what data or model you need. The ring begins at Problem.

How do I know if I have enough data?

Data understanding: look at counts, charts and gaps before modelling.

Why remove rows? Is that not throwing data away?

Wrong or duplicate rows teach the model wrong patterns. A smaller clean set beats a bigger dirty one.

Why hide the test data?

Like an exam with unseen questions, it shows whether the model learnt the pattern or just memorised.

Is the project finished once the model is deployed?

No. Feedback shows when the model gets worse, and the cycle restarts.

What split should I use?

Commonly 70:30 or 80:20. Try the slider and see the trade-off.

What is data science methodology?

Data science uses data to answer questions and make predictions. A methodology is a fixed set of steps that guides the work, so nothing important is missed. The common data science methodology has 10 stages arranged in a cycle. It is the backbone of an AI capstone project.

  1. Business understanding (the problem)
  2. Analytic approach
  3. Data requirements
  4. Data collection
  5. Data understanding
  6. Data preparation
  7. Modelling
  8. Evaluation
  9. Deployment
  10. Feedback

The stages can loop back. For example, data understanding may show we need more data, so we return to collection.

From problem to approach

1. Business understanding

Write the problem as one clear question. Who will use the answer? How will we know we have succeeded? A tool such as the 5W1H questions (who, what, where, when, why, how) helps.

2. Analytic approach

Choose the kind of answer:

Data requirements, collection, understanding and preparation

3. Data requirements

List the data needed: which features (columns), how many rows, what format and from what time.

4. Data collection

Get data from records, sensors, surveys, websites or open datasets. Get permission and protect personal data.

5. Data understanding

Use summaries (mean, min, max), charts and correlations to see if the data is complete and sensible.

6. Data preparation

Usually the longest stage. Remove duplicates, handle missing values (fill with a mean, or drop the row), fix impossible values, convert text to numbers and create new useful features (feature engineering).

Modelling, validation and evaluation

7. Modelling

Train a model, such as a decision tree or linear regression. To test it fairly we must hide some data from it.

8. Evaluation

For classification, use accuracy = correct predictions ÷ all predictions (and a confusion matrix). For regression, use errors:

MSE (mean squared error) = average of (actual − predicted)²

RMSE = √MSE, in the same units as the data. Smaller is better.

Example: actual 3, 5, 2; predicted 2, 5, 4. Errors 1, 0, −2; squares 1, 0, 4. MSE = 5 ÷ 3 ≈ 1.67; RMSE ≈ 1.29.

Deployment, feedback and the capstone project

9. Deployment

Put the model to work: in an app, a website or a report that people use.

10. Feedback

Collect results from real use. If the world changes (new menu, new prices), the model gets worse, so we retrain it. Feedback makes data science a cycle, not a straight line.

Using it in a capstone project

A capstone project follows the same stages: choose a real problem, state the approach, gather and clean data, build and test a model, show results, and explain limits and ethics (bias, privacy).

Try it: the 3D and at home

In the free-play step, press each stage and say aloud what you would do for a school canteen. Move the split slider: with 50% training the model learns from less data; with 90% the test is very small. Why is 70 to 80% a common choice?

At home: plan a mini project, "Which day of the week does our family use the most milk?" Write the 10 stages on paper, collect data for two weeks, then check your prediction.

Key formulas and definitions

Worked examples

1. A bank wants to know if a loan applicant will repay or not. Which analytic approach fits?

The answer is a category (repay / not repay), so it is classification.

2. A dataset has 1,000 rows. With an 80:20 split, how many rows are used for training and testing?

Training = 0.8 × 1,000 = 800 rows. Testing = 200 rows.

3. Actual values 10, 12, 8; predicted 11, 12, 6. Find MSE and RMSE.

Errors −1, 0, 2. Squares 1, 0, 4. MSE = 5 ÷ 3 ≈ 1.67. RMSE = √1.67 ≈ 1.29.

4. In 5-fold cross-validation on 500 rows, how many rows are tested in each round, and how many rounds are there?

500 ÷ 5 = 100 rows per fold are tested each round, and there are 5 rounds.

5. A model predicts 45 out of 50 test emails correctly. Find its accuracy.

Accuracy = 45 ÷ 50 = 0.9 = 90%.

6. A traffic model worked well last year but now gives poor results after a new metro line opened. Which stage fixes this?

Feedback: collect new data, go back through preparation and modelling, and retrain the model.

Common mistakes

Practice quiz

1. The first stage of data science methodology is:
2. Predicting tomorrow's temperature as a number is:
3. Which stage usually takes the most time?
4. RMSE is:
5. Why keep test data hidden from the model?

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What are the steps of data science methodology?

Business understanding, analytic approach, data requirements, data collection, data understanding, data preparation, modelling, evaluation, deployment and feedback.

What is the difference between MSE and RMSE?

MSE is the average of squared errors. RMSE is its square root, so it is in the same units as the data and easier to read.

What is cross-validation?

A way to test a model by splitting data into k folds and letting each fold be the test set once, then averaging the scores.

Where this is taught

CBSE (India)Class 12Data Science Methodology

Learn first

Learn next

Related lessons

All Computer Science lessons