What are data and why do geographers need them?
Data are numbers or facts we record, such as rainfall, crop output or the number of people in a village. Geographers use data to compare places, find patterns and plan (roads, schools, water).
Data can be about quantity (how much, how many) or quality (type of soil, type of house).
Sources of data: primary, secondary and unpublished
Primary sources
Data you collect yourself, first-hand. Methods: personal observation (look and record), interview (talk to people), questionnaire or schedule (a printed list of questions), and other methods like measuring soil or water in the field.
Secondary sources
Data someone else collected and published. Examples: Census of India, National Sample Survey, weather reports, government yearbooks, maps, books, journals, newspapers, international reports (UN, World Bank) and official websites.
Unpublished sources
Records kept by offices but not printed for the public: village land records (patwari), municipal files, school registers, hospital records, company files.
Which is better?
Primary data fit your exact question but take time and money. Secondary data are quick and cover large areas, but may be old or collected for another purpose.
Tabulation: putting data into a table
Tabulation means arranging data in rows and columns. A good table has a number, a title, column headings, units (cm, km², %), the body with the values, and the source at the bottom.
Tables make big data easy to read and compare. You can also add totals, averages and percentages. Absolute numbers, ratios and percentages can all be tabulated.
Grouping data into classes
When there are many values, we group them into classes (like 0–10, 10–20). Each class has a lower limit and an upper limit. The difference is the class interval (class width).
Steps
- Find the range = highest − lowest value.
- Choose the number of classes (usually 5 to 15) and a round class width.
- Put a tally mark for each value in its class (every fifth mark crosses the four: ||||).
- Count the tallies. This is the frequency (f). The frequencies must add up to the total number of values (N).
Exclusive and inclusive methods
Exclusive: 0–10, 10–20… the upper limit is not included, so 10 goes to 10–20. Inclusive: 0–9, 10–19… both limits are included. Cumulative frequency is the running total of frequencies.
Frequency polygon
A frequency polygon is a line graph of grouped data. Steps:
- Find the midpoint of each class: (lower + upper) ÷ 2.
- Plot midpoint on the x-axis and frequency on the y-axis.
- Add one extra class with frequency 0 at each end, so the shape closes on the x-axis.
- Join the points with straight lines.
You can draw it on top of a histogram (joining the midpoints of the bar tops) or alone. Two polygons on one graph help compare two sets of data. An ogive is a similar line for cumulative frequency.
Key formulas and definitions
- Range = highest value − lowest value
- Class interval (width) = upper limit − lower limit
- Midpoint = (lower limit + upper limit) ÷ 2
- Σf = N (frequencies add up to the total)
- Cumulative frequency = running total of f
Worked examples
1. You count the vehicles passing your school gate for one hour. Is this primary or secondary data?
Primary. You collected it yourself by observation.
2. Data taken from the Census of India report to compare literacy of two districts: what type?
Secondary data, because it was collected and published by the Census office.
3. Values range from 12 to 87. You want 8 classes. What class width should you choose?
Range = 87 − 12 = 75. 75 ÷ 8 ≈ 9.4, so round up to 10. Classes: 10–20, 20–30, …, 80–90 (8 classes).
4. Find the midpoints of classes 20–30, 30–40 and 40–50.
(20+30)/2 = 25, (30+40)/2 = 35, (40+50)/2 = 45.
5. Rainfall values: 12, 25, 7, 31, 18, 22, 9, 27, 15, 34, 21, 13, 28, 19, 5, 24, 16, 38, 23, 11 (cm). Group them in classes of 10 (exclusive).
0–10: 7, 9, 5 → f = 3. 10–20: 12, 18, 15, 13, 19, 16, 11 → f = 7. 20–30: 25, 22, 27, 21, 28, 24, 23 → f = 7. 30–40: 31, 34, 38 → f = 3. Check: 3+7+7+3 = 20 = N.
6. Using the table above, list the points you plot for the frequency polygon.
Add empty classes −10–0 and 40–50 with f = 0. Points (midpoint, f): (−5, 0), (5, 3), (15, 7), (25, 7), (35, 3), (45, 0). Join them with straight lines.
7. Convert the inclusive classes 10–19, 20–29, 30–39 to exclusive classes.
Gap between 19 and 20 is 1; half is 0.5. Subtract 0.5 from lower limits and add 0.5 to upper limits: 9.5–19.5, 19.5–29.5, 29.5–39.5.
Common mistakes
- Calling Census data primary. It is secondary for you because someone else collected and published it.
- Putting a value on the boundary (like 20) in the lower class in the exclusive method. It belongs to 20–30.
- Forgetting to add the zero classes at both ends of a frequency polygon, so it does not touch the x-axis.
- Making a table without a title, units or source.