Why collect data?
Data means facts in numbers that help us answer a question. Example question: "How much rice do families in our town eat each month?" To answer it we must first collect data. The person who collects data is the investigator. The person who gives the answers is the respondent. The people or things we study make up the population (also called the universe).
Primary and secondary sources
- Primary data: collected for the first time by the investigator for their own purpose. It is original and fits the need well, but takes more time, money and effort.
- Secondary data: already collected and published by someone else. It is cheap and quick, but may not fit your purpose exactly, so check it before use.
Sources of secondary data
Government reports (Census of India, NSSO reports, Economic Survey), Reserve Bank of India reports, newspapers, research papers and websites of official bodies.
Methods of collecting primary data
1. Personal interview
The investigator meets the respondent face to face. Answers are complete and doubts can be cleared on the spot. But it is costly and slow for a large area, and the investigator may influence the answers.
2. Mailing questionnaire
A printed or online form is sent; people fill it and send it back. It is cheap and reaches far-off places. But many people do not reply, and it works only with literate people.
3. Telephone interview
Questions are asked on the phone. It is quick and cheap, but people without phones are left out and you cannot see reactions.
Preparing a good questionnaire
- Keep it short. Ask only what is needed.
- Use simple, clear words; one idea per question.
- Move from general to specific questions.
- Avoid leading questions, like "Don't you think prices are too high?"
- Avoid questions that hurt feelings or need calculations.
- Prefer closed questions with options (yes/no or multiple choice) when possible.
Pilot survey
A small trial survey done before the main one. It shows weak questions, time needed and cost.
Census method and sample method
- Census (complete enumeration): every member of the population is studied. It is thorough, but costly and slow.
- Sample method: a small group, chosen to represent the population, is studied. It saves time and money and often gives reliable results. A good sample is representative: it looks like the population in the ways that matter.
Random and non-random sampling
Random sampling
Every member has an equal chance of being picked. Methods: lottery (names on slips in a drum) or a table of random numbers (today, a computer).
Non-random sampling
Not everyone has an equal chance. The investigator picks by convenience (people nearby), judgement (whom they think is typical) or quota (a fixed number from each group). It is easier but can be biased.
Sampling and non-sampling errors
Sampling error = value from the sample − true value of the population. It happens because we studied only a part. A bigger, well-chosen sample makes it smaller.
Non-sampling errors can happen even in a census:
- Errors in data acquisition: wrong measuring, wrong copying, confusing questions.
- Non-response errors: some people do not answer.
- Sampling bias: some members of the population could never be chosen.
Census of India and NSSO
Census of India
It counts every person in India, once every 10 years. The first full, regular census was in 1881 and the first after independence in 1951. It collects data on population size, age, sex, literacy, occupation, housing and migration. It is run by the Office of the Registrar General and Census Commissioner.
NSSO (National Sample Survey Office)
Set up in 1950 to run large sample surveys across India. In 2019 it was merged with the Central Statistical Office into the National Statistical Office (NSO). Its survey rounds study household spending, jobs and unemployment, literacy, health and land holdings. Results appear in reports and in the journal Sarvekshana.
Try it: a mini survey at home
Write 3 short, clear questions, such as "How many litres of milk does your family buy each week?" Ask 10 neighbours (primary data). Now pick 3 of them by lottery using paper slips (a random sample). Compare the average of the 3 with the average of all 10. The difference is your sampling error. Then try the last 3D step with sizes 4, 10 and 30.
Board exam pattern
Common questions: difference between primary and secondary data (3 marks), census vs sample (3 marks), qualities of a good questionnaire (4 marks), sampling vs non-sampling errors, and short notes on the Census of India and NSSO.
Key formulas and definitions
- Primary data: first-hand, collected by the investigator
- Secondary data: already collected by someone else
- Census: every unit of the population is studied
- Sample: a representative part of the population is studied
- Sampling error = sample value − population value
- Random sample: every unit has an equal chance of being picked
Worked examples
1. A student asks 50 shopkeepers about their daily sales. Is this primary or secondary data?
Primary. The student collected it first-hand for her own study.
2. Incomes of a population of 5 farmers are ₹5,000, ₹6,000, ₹7,000, ₹8,000 and ₹9,000. A sample picks the 2nd and 5th farmers. Find the sampling error of the mean.
Population mean = 35,000 ÷ 5 = ₹7,000. Sample mean = (6,000 + 9,000) ÷ 2 = ₹7,500. Sampling error = 7,500 − 7,000 = ₹500.
3. Find what is wrong with the question: "Don't you agree that our school canteen is too costly?"
It is a leading question: it pushes the respondent towards "yes". Better: "How do you find canteen prices? (low / fair / high)".
4. The government wants to know the literacy rate of every village in India. Should it use census or sample? Why?
Census. It needs a figure for each village, so every unit must be counted. A sample would miss many villages.
Common mistakes
- Calling Census data "primary" when you just read it in a report. For you it is secondary.
- Thinking a sample is always worse than a census. A good random sample is cheaper and can be very accurate.
- Saying non-sampling errors happen only in samples. They can occur in a census too.
- Writing leading or double questions, like "Do you like tea and coffee?"