What is data science methodology?
Data science uses data to answer questions and make predictions. A methodology is a fixed set of steps that guides the work, so nothing important is missed. The common data science methodology has 10 stages arranged in a cycle. It is the backbone of an AI capstone project.
- Business understanding (the problem)
- Analytic approach
- Data requirements
- Data collection
- Data understanding
- Data preparation
- Modelling
- Evaluation
- Deployment
- Feedback
The stages can loop back. For example, data understanding may show we need more data, so we return to collection.
From problem to approach
1. Business understanding
Write the problem as one clear question. Who will use the answer? How will we know we have succeeded? A tool such as the 5W1H questions (who, what, where, when, why, how) helps.
2. Analytic approach
Choose the kind of answer:
- Classification: a category (spam or not spam).
- Regression: a number (tomorrow's sales).
- Clustering: natural groups with no labels (types of customers).
- Descriptive: what happened (charts and averages).
Data requirements, collection, understanding and preparation
3. Data requirements
List the data needed: which features (columns), how many rows, what format and from what time.
4. Data collection
Get data from records, sensors, surveys, websites or open datasets. Get permission and protect personal data.
5. Data understanding
Use summaries (mean, min, max), charts and correlations to see if the data is complete and sensible.
6. Data preparation
Usually the longest stage. Remove duplicates, handle missing values (fill with a mean, or drop the row), fix impossible values, convert text to numbers and create new useful features (feature engineering).
Modelling, validation and evaluation
7. Modelling
Train a model, such as a decision tree or linear regression. To test it fairly we must hide some data from it.
- Train-test split: often 80% for training and 20% for testing.
- k-fold cross-validation: split the data into k equal parts. Train on k − 1 parts and test on the one left, k times, each part taking a turn as the test. Average the k scores. It gives a more reliable score when data is small.
8. Evaluation
For classification, use accuracy = correct predictions ÷ all predictions (and a confusion matrix). For regression, use errors:
MSE (mean squared error) = average of (actual − predicted)²
RMSE = √MSE, in the same units as the data. Smaller is better.
Example: actual 3, 5, 2; predicted 2, 5, 4. Errors 1, 0, −2; squares 1, 0, 4. MSE = 5 ÷ 3 ≈ 1.67; RMSE ≈ 1.29.
Deployment, feedback and the capstone project
9. Deployment
Put the model to work: in an app, a website or a report that people use.
10. Feedback
Collect results from real use. If the world changes (new menu, new prices), the model gets worse, so we retrain it. Feedback makes data science a cycle, not a straight line.
Using it in a capstone project
A capstone project follows the same stages: choose a real problem, state the approach, gather and clean data, build and test a model, show results, and explain limits and ethics (bias, privacy).
Try it: the 3D and at home
In the free-play step, press each stage and say aloud what you would do for a school canteen. Move the split slider: with 50% training the model learns from less data; with 90% the test is very small. Why is 70 to 80% a common choice?
At home: plan a mini project, "Which day of the week does our family use the most milk?" Write the 10 stages on paper, collect data for two weeks, then check your prediction.
Key formulas and definitions
- Accuracy = correct predictions ÷ total predictions
- Error = actual − predicted
- MSE = (1/n) Σ (actual − predicted)²
- RMSE = √MSE
- k-fold cross-validation: k rounds; each round trains on k − 1 parts and tests on 1 part; final score = average
Worked examples
1. A bank wants to know if a loan applicant will repay or not. Which analytic approach fits?
The answer is a category (repay / not repay), so it is classification.
2. A dataset has 1,000 rows. With an 80:20 split, how many rows are used for training and testing?
Training = 0.8 × 1,000 = 800 rows. Testing = 200 rows.
3. Actual values 10, 12, 8; predicted 11, 12, 6. Find MSE and RMSE.
Errors −1, 0, 2. Squares 1, 0, 4. MSE = 5 ÷ 3 ≈ 1.67. RMSE = √1.67 ≈ 1.29.
4. In 5-fold cross-validation on 500 rows, how many rows are tested in each round, and how many rounds are there?
500 ÷ 5 = 100 rows per fold are tested each round, and there are 5 rounds.
5. A model predicts 45 out of 50 test emails correctly. Find its accuracy.
Accuracy = 45 ÷ 50 = 0.9 = 90%.
6. A traffic model worked well last year but now gives poor results after a new metro line opened. Which stage fixes this?
Feedback: collect new data, go back through preparation and modelling, and retrain the model.
Common mistakes
- Testing the model on the same data it was trained on. The score looks great but means nothing; use unseen test data.
- Jumping to modelling without a clear problem. A precise question decides the data and the approach.
- Skipping data cleaning. Duplicates, missing and wrong values make the model learn wrong patterns.
- Thinking the work ends at deployment. Models need feedback and retraining as the world changes.