๐Ÿ“˜ CodingMarble Learn

Pandas Series: One Labelled Column of Data

Pandas is a Python library for working with data in tables, and Matplotlib is a library for drawing charts. A Series is one column of values where every value has a label called its index. You can make a Series from a list or NumPy array, from a dictionary or from a single scalar value. Maths between two Series matches labels, and a label present in only one gives NaN. head() and tail() show the first or last rows, and you can pick values by label, by position, by slice or by a condition.

๐ŸŽฌ Step-by-step story

  1. Pandas handles tables of data; Matplotlib draws charts.
  2. A Series is one column: each value has a label called the index.
  3. Make a Series from an array, a dictionary or one repeated value.
  4. Maths on two Series matches labels; a missing label gives NaN.
  5. head, tail, labels, positions, slices and conditions pick values.
  6. Your turn: pick an operation and see which values light up.

Tip: drag the 3D scene to turn it. Use two fingers to zoom.

๐Ÿค” Common doubts, cleared

What is the difference between a Series and a Python list?

A list has only positions. A Series has values plus labels (the index), and maths works on all values at once, matched by label.

Why did my dictionary keys become the index?

Pandas reads a dictionary as label โ†’ value. So keys go to the index and values go to the data.

Why do I get NaN when I add two Series?

Pandas adds values that share the same label. A label that is in only one Series has no partner, so the answer is NaN. Use add() with fill_value=0 to avoid it.

Why does s[1:3] leave out position 3 but s['a':'c'] includes 'c'?

Position slices follow normal Python rules (end left out). Label slices include the end because labels have no 'one before' to stop at.

What is the difference between loc and iloc?

loc uses labels (s.loc['Zoya']). iloc uses positions counting from 0 (s.iloc[3]).

How do I check my guess for a slice quickly?

Use step 6: press the slice button and see which bars rise.

Pandas and Matplotlib libraries

A library is a ready-made box of code that we can use in our own program. Python has many libraries for data.

pip install pandas matplotlib     # once, in the command prompt
import pandas as pd               # pd is a short name (alias)
import numpy as np
import matplotlib.pyplot as plt

What is a Series?

A Series is a one-dimensional labelled array. That means: one column of values, and every value has a label. The labels together are called the index. The values are usually all of the same type.

s = pd.Series([92, 75, 58], index=['Asha', 'Ravi', 'Mohan'])
print(s)
Asha     92
Ravi     75
Mohan    58
dtype: int64

If you do not give an index, pandas uses 0, 1, 2, โ€ฆ by itself. Useful attributes: s.index, s.values, s.dtype, s.size, s.shape, s.empty, s.name.

Creating a Series from ndarray, dictionary and scalar

From a NumPy array (ndarray)

a = np.array([10, 20, 30])
s = pd.Series(a)                         # index 0, 1, 2
s = pd.Series(a, index=['x', 'y', 'z'])  # index length must match

From a dictionary

d = {'Jan': 31, 'Feb': 28, 'Mar': 31}
s = pd.Series(d)       # keys become the index, values become the values

From a scalar (one value)

s = pd.Series(5, index=['a', 'b', 'c', 'd'])   # 5 is repeated 4 times

For a scalar you must give an index, otherwise you get a Series with just one value.

Maths operations on Series

With one number, the operation works on every value: s * 2, s + 5, s ** 2.

With two Series, pandas matches values by index label, not by position. If a label is found in only one of them, the answer for that label is NaN (Not a Number, meaning missing).

s1 = pd.Series([10, 20, 30], index=['a', 'b', 'c'])
s2 = pd.Series([1, 2, 4],    index=['a', 'b', 'd'])
s1 + s2          # a 11.0, b 22.0, c NaN, d NaN
s1.add(s2, fill_value=0)   # a 11, b 22, c 30, d 4

Methods: add(), sub(), mul(), div(). They accept fill_value, which the plain operators do not. Because NaN is a float, the result type becomes float.

head() and tail()

s.head(n) returns the first n values; s.tail(n) returns the last n. With no number, both use 5. They are handy to peek at a long Series without printing all of it.

s.head(2)   # Asha 92, Ravi 75
s.tail(3)   # Mohan 58, Zoya 81, Kabir 66
s.head()    # first 5

Selection, indexing and slicing

Indexing (one value)

s['Zoya']      # by label โ†’ 81
s.loc['Zoya']  # by label (clear way)
s.iloc[3]      # by position โ†’ 81
s[['Asha', 'Kabir']]   # a list of labels โ†’ a smaller Series

Slicing (a range)

s[1:4]            # positions 1, 2, 3 (end NOT included)
s['Ravi':'Zoya']  # labels Ravi to Zoya (end IS included)
s[::-1]           # reverse order

Selection by condition (boolean indexing)

s[s > 70]      # only values above 70
s[s == 58]

You can also change values: s['Mohan'] = 60 or s[1:3] = 0.

Try it: your own marks Series

Open Python (IDLE, Jupyter or any online Python). Make a dictionary of your 5 subjects and marks, then turn it into a Series. Before running each line, write your guess on paper: s.head(2), s[s > 80], s[1:3], s + 5. Run them and tick the guesses you got right. Then try the same in step 6 of the 3D.

Key formulas and definitions

Worked examples

1. Create a Series of the days in the first three months using a dictionary, and print it.

d = {'Jan': 31, 'Feb': 28, 'Mar': 31}; s = pd.Series(d); print(s) โ†’ Jan 31, Feb 28, Mar 31, dtype: int64. The keys became the index.

2. Create a Series of four 100s with index 'w', 'x', 'y', 'z'.

s = pd.Series(100, index=['w', 'x', 'y', 'z']). The scalar 100 is repeated for each label.

3. s = pd.Series([92, 75, 58, 81, 66], index=['Asha','Ravi','Mohan','Zoya','Kabir']). What do s[1:4] and s['Ravi':'Zoya'] give?

Both give Ravi 75, Mohan 58, Zoya 81. Position slice 1:4 takes positions 1, 2, 3 (4 left out). Label slice includes the end label Zoya.

4. For the same s, write code to show only students with marks above 70, and count them.

s[s > 70] โ†’ Asha 92, Ravi 75, Zoya 81. Count: s[s > 70].size or len(s[s > 70]) โ†’ 3.

5. s1 = pd.Series([10, 20, 30], index=['a','b','c']); s2 = pd.Series([1, 2, 4], index=['a','b','d']). Find s1 + s2 and s1.add(s2, fill_value=0).

s1 + s2 โ†’ a 11.0, b 22.0, c NaN, d NaN (c and d have no partner). With fill_value=0: a 11, b 22, c 30, d 4.

6. Using the marks Series s, give 5 grace marks to everyone below 70 in one line.

s[s < 70] = s[s < 70] + 5 โ†’ Mohan 58 โ†’ 63, Kabir 66 โ†’ 71; others unchanged.

Common mistakes

Practice quiz

1. A Series is best described as:
2. When a Series is made from a dictionary, the keys become the:
3. s.head() with no argument returns how many values?
4. If a label is in s1 but not in s2, s1 + s2 gives for that label:
5. Which gives the value at position 3?

Practice: answer these yourself

Type or choose your answer, then press Check. Use a hint if you are stuck; the full solution appears after you answer.

Frequently asked questions

What is a Series in pandas Class 12?

A Series is a one-dimensional labelled array: one column of values, each with an index label.

How many ways can we create a Series?

From a list or NumPy ndarray, from a dictionary (keys become index) and from a scalar value with an index.

What do head() and tail() return?

head(n) returns the first n values and tail(n) the last n; both use 5 when n is not given.

Where this is taught

CBSE (India)Class 12Data Handling using Pandas and Data Visualization

Learn first

Learn next

Related lessons

All Informatics Practices lessons