Raw data and why we organise it
Raw data is data just as it was collected, with no order. It is hard to read. Classification means putting data into groups (classes) of similar items. It makes data short, clear and easy to compare.
Kinds of classification
- Chronological: by time (population of India in 1951, 1961, 1971 …).
- Spatial: by place (wheat output of different states).
- Qualitative: by a quality that cannot be measured (literate or not, male or female).
- Quantitative: by a measurable quantity (marks, income, height).
Types of variables
A variable is a quantity whose value changes from one item to another, like age or income.
- Discrete variable: takes only certain separate values, usually whole numbers. Example: number of children in a family (1, 2, 3, never 2.5).
- Continuous variable: can take any value in a range, including fractions. Example: height (150 cm, 150.4 cm, 150.45 cm …), weight, time.
Frequency distribution
A frequency distribution is a table that shows classes and the number of values in each class. Frequency (f) = how many times a value or class occurs.
Words to know
- Class: a group, like 30–40.
- Class limits: the two end values. 30 is the lower limit, 40 the upper limit.
- Class interval (width) = upper limit − lower limit.
- Class mark (mid-value) = (lower limit + upper limit) ÷ 2.
- Tally marks: a stroke for each value; every fifth stroke is drawn across the first four.
How to make one
- Find the range: largest − smallest.
- Decide the number of classes (often 5 to 15) and the width.
- Write the classes, go through the data and put a tally for each value.
- Count the tallies to get f. Check that Σf equals the number of values.
Exclusive and inclusive methods
- Exclusive method: classes 0–10, 10–20, 20–30. The upper limit is excluded: a value of 10 goes into 10–20. Used for continuous data.
- Inclusive method: classes 0–9, 10–19, 20–29. Both limits are included. Used for discrete data.
Adjusting inclusive classes
Find the gap: lower limit of next class − upper limit of this class (10 − 9 = 1). Half of it is 0.5. Subtract 0.5 from every lower limit and add 0.5 to every upper limit: 0–9 becomes −0.5–9.5, 10–19 becomes 9.5–19.5.
Equal and unequal classes, loss of information
Usually all classes have equal width. Sometimes, when data is very spread out (like incomes), unequal classes are used. When we group data, we treat all values in a class as if they were the class mark. So some detail is lost. This is called loss of information. Wider classes lose more.
Frequency array and bivariate distribution
A frequency array is used for a discrete variable: each value is written with its frequency (family size 1: 5 families, 2: 15, 3: 25 …).
A bivariate frequency distribution shows two variables at once in one table, for example sales of 20 companies in rows and their advertising spending in columns. Each cell tells how many companies fall in both classes.
Try it: sort your family's ages
Write the ages of 15 relatives or neighbours. Make classes 0–10, 10–20, 20–30 and so on. Put a tally for each age, then count. Which class has the highest frequency? Now try classes of width 20. Which table tells you more? Check the same idea with the width picker on the last 3D step.
Board exam pattern
Expect questions like: make a frequency distribution with given classes (4 marks), convert inclusive classes into exclusive, find class marks, and difference between discrete and continuous variables. Always show the tally column and check Σf.
Key formulas and definitions
- Range = largest value − smallest value
- Class interval (width) = upper limit − lower limit
- Class mark = (lower limit + upper limit) ÷ 2
- Number of classes ≈ range ÷ width
- Adjustment for inclusive classes = (next lower limit − this upper limit) ÷ 2
- Σf = total number of observations
Worked examples
1. Find the class mark and width of the class 40–60.
Class mark = (40 + 60) ÷ 2 = 50. Width = 60 − 40 = 20.
2. Is "number of cars sold in a day" discrete or continuous? And "petrol used in a day"?
Cars sold is discrete (whole cars only). Petrol used is continuous (it can be 12.35 litres).
3. Convert 10–19, 20–29, 30–39 into exclusive classes.
Gap = 20 − 19 = 1, adjustment = 0.5. New classes: 9.5–19.5, 19.5–29.5, 29.5–39.5.
4. Make a frequency table with classes 0–10, 10–20, 20–30 for: 5, 12, 18, 25, 10, 3, 27, 20, 14, 8.
0–10: 5, 3, 8 → f = 3. 10–20: 12, 18, 10, 14 → f = 4 (10 goes here). 20–30: 25, 27, 20 → f = 3. Σf = 10.
5. The class marks of a table are 15, 25, 35, 45. Find the classes.
Width = 25 − 15 = 10. Half width = 5. Classes: 10–20, 20–30, 30–40, 40–50.
6. Marks range from 12 to 72. If you want 7 classes, what width should you use, and where should the first class start?
Range = 72 − 12 = 60. Width ≈ 60 ÷ 7 ≈ 8.6, round up to 10 for easy classes. Start at 10: 10–20, 20–30, … 70–80 covers all values (7 classes).
Common mistakes
- Putting the value 20 in class 10–20 in the exclusive method. It goes in 20–30.
- Forgetting to adjust inclusive classes before finding the median or drawing a histogram.
- Calling "number of students" continuous. You cannot have 30.5 students; it is discrete.
- Not checking that Σf equals the total number of values.