📘 CodingMarble Learn

Collection of Data: Primary and Secondary Sources, Sampling, Census of India and NSSO

Data can be primary (collected first-hand by you) or secondary (already collected by someone else). Primary data comes from personal interviews, mailed questionnaires or telephone interviews. We may study everyone (census) or a small part (sample). Samples can be random or non-random. Results can have sampling errors and non-sampling errors. In India, the Census counts everyone every 10 years, and the NSSO (now NSO) runs sample surveys.

🎬 Step-by-step story

  1. The investigator walks to each house and asks questions: that is primary data. Picking a figure from a printed report is secondary data.
  2. There are three ways to ask: face to face, by a mailed questionnaire, or by phone. Each costs a different amount of time and money.
  3. Census means asking every one of the 40 people. Sample means asking only a few, say 8. The Census of India counts everyone; the NSSO surveys samples.
  4. In a random sample, a lottery gives everyone an equal chance. In a non-random sample, the investigator picks people by choice, which can bring bias.
  5. The front-row sample average is far from the true average. That gap is the sampling error. Wrong measuring or no reply cause non-sampling errors.
  6. Free play: change the sample size and draw again. See the error shrink as the sample grows.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

🤔 Common doubts, cleared

Is Census data primary or secondary?

For the Census office it is primary. For you, reading it in a report, it is secondary.

Which method is best for villagers who cannot read?

Personal interview, because a mailed form needs reading and writing.

Why not always use a census?

It is costly and slow. A good sample gives nearly the same answer at a fraction of the cost.

What makes a sample random?

Every unit has the same chance of being picked, like names drawn by lottery.

Can we remove sampling error completely?

Only by studying everyone (census). A bigger random sample makes it small.

What is the difference between sampling error and non-sampling error?

Sampling error comes from studying only a part. Non-sampling errors come from mistakes in asking, measuring or recording, or from people not replying.

Why collect data?

Data means facts in numbers that help us answer a question. Example question: "How much rice do families in our town eat each month?" To answer it we must first collect data. The person who collects data is the investigator. The person who gives the answers is the respondent. The people or things we study make up the population (also called the universe).

Primary and secondary sources

Sources of secondary data

Government reports (Census of India, NSSO reports, Economic Survey), Reserve Bank of India reports, newspapers, research papers and websites of official bodies.

Methods of collecting primary data

1. Personal interview

The investigator meets the respondent face to face. Answers are complete and doubts can be cleared on the spot. But it is costly and slow for a large area, and the investigator may influence the answers.

2. Mailing questionnaire

A printed or online form is sent; people fill it and send it back. It is cheap and reaches far-off places. But many people do not reply, and it works only with literate people.

3. Telephone interview

Questions are asked on the phone. It is quick and cheap, but people without phones are left out and you cannot see reactions.

Preparing a good questionnaire

Pilot survey

A small trial survey done before the main one. It shows weak questions, time needed and cost.

Census method and sample method

Random and non-random sampling

Random sampling

Every member has an equal chance of being picked. Methods: lottery (names on slips in a drum) or a table of random numbers (today, a computer).

Non-random sampling

Not everyone has an equal chance. The investigator picks by convenience (people nearby), judgement (whom they think is typical) or quota (a fixed number from each group). It is easier but can be biased.

Sampling and non-sampling errors

Sampling error = value from the sample − true value of the population. It happens because we studied only a part. A bigger, well-chosen sample makes it smaller.

Non-sampling errors can happen even in a census:

Census of India and NSSO

Census of India

It counts every person in India, once every 10 years. The first full, regular census was in 1881 and the first after independence in 1951. It collects data on population size, age, sex, literacy, occupation, housing and migration. It is run by the Office of the Registrar General and Census Commissioner.

NSSO (National Sample Survey Office)

Set up in 1950 to run large sample surveys across India. In 2019 it was merged with the Central Statistical Office into the National Statistical Office (NSO). Its survey rounds study household spending, jobs and unemployment, literacy, health and land holdings. Results appear in reports and in the journal Sarvekshana.

Try it: a mini survey at home

Write 3 short, clear questions, such as "How many litres of milk does your family buy each week?" Ask 10 neighbours (primary data). Now pick 3 of them by lottery using paper slips (a random sample). Compare the average of the 3 with the average of all 10. The difference is your sampling error. Then try the last 3D step with sizes 4, 10 and 30.

Board exam pattern

Common questions: difference between primary and secondary data (3 marks), census vs sample (3 marks), qualities of a good questionnaire (4 marks), sampling vs non-sampling errors, and short notes on the Census of India and NSSO.

Key formulas and definitions

Worked examples

1. A student asks 50 shopkeepers about their daily sales. Is this primary or secondary data?

Primary. The student collected it first-hand for her own study.

2. Incomes of a population of 5 farmers are ₹5,000, ₹6,000, ₹7,000, ₹8,000 and ₹9,000. A sample picks the 2nd and 5th farmers. Find the sampling error of the mean.

Population mean = 35,000 ÷ 5 = ₹7,000. Sample mean = (6,000 + 9,000) ÷ 2 = ₹7,500. Sampling error = 7,500 − 7,000 = ₹500.

3. Find what is wrong with the question: "Don't you agree that our school canteen is too costly?"

It is a leading question: it pushes the respondent towards "yes". Better: "How do you find canteen prices? (low / fair / high)".

4. The government wants to know the literacy rate of every village in India. Should it use census or sample? Why?

Census. It needs a figure for each village, so every unit must be counted. A sample would miss many villages.

Common mistakes

Practice quiz

1. Data collected first-hand by the investigator is called:
2. In a random sample, each unit has:
3. The Census of India is held every:
4. A small trial survey before the main survey is a:
5. Which is a non-sampling error?

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is the difference between primary and secondary data?

Primary data is collected first-hand for your purpose; secondary data was already collected by someone else and is reused.

What is NSSO?

The National Sample Survey Office, set up in 1950, runs nationwide sample surveys on spending, jobs, health and more. Since 2019 it is part of the NSO.

What is a sampling error?

The difference between a value found from a sample and the true value of the whole population.

Where this is taught

Canada (Ontario)Grade 12C. Organization of Data for Analysis
Spain2º ESOScientific project
Spain3º ESOScientific project
Spain4º ESOScientific project
CBSE (India)Class 11Collection, Organisation and Presentation of Data
England (GCSE, A level)Year 101. The collection of data

Learn first

Learn next

Related lessons

All Economics lessons