Pandas and Matplotlib libraries
A library is a ready-made box of code that we can use in our own program. Python has many libraries for data.
- NumPy: fast arrays of numbers (you met it in Class 11).
- Pandas: data in rows and columns, like a spreadsheet. Its two main data structures are Series (one column) and DataFrame (a full table).
- Matplotlib: draws charts such as line plots, bar graphs and histograms.
pip install pandas matplotlib # once, in the command prompt import pandas as pd # pd is a short name (alias) import numpy as np import matplotlib.pyplot as plt
What is a Series?
A Series is a one-dimensional labelled array. That means: one column of values, and every value has a label. The labels together are called the index. The values are usually all of the same type.
s = pd.Series([92, 75, 58], index=['Asha', 'Ravi', 'Mohan']) print(s) Asha 92 Ravi 75 Mohan 58 dtype: int64
If you do not give an index, pandas uses 0, 1, 2, โฆ by itself. Useful attributes: s.index, s.values, s.dtype, s.size, s.shape, s.empty, s.name.
Creating a Series from ndarray, dictionary and scalar
From a NumPy array (ndarray)
a = np.array([10, 20, 30]) s = pd.Series(a) # index 0, 1, 2 s = pd.Series(a, index=['x', 'y', 'z']) # index length must match
From a dictionary
d = {'Jan': 31, 'Feb': 28, 'Mar': 31}
s = pd.Series(d) # keys become the index, values become the valuesFrom a scalar (one value)
s = pd.Series(5, index=['a', 'b', 'c', 'd']) # 5 is repeated 4 times
For a scalar you must give an index, otherwise you get a Series with just one value.
Maths operations on Series
With one number, the operation works on every value: s * 2, s + 5, s ** 2.
With two Series, pandas matches values by index label, not by position. If a label is found in only one of them, the answer for that label is NaN (Not a Number, meaning missing).
s1 = pd.Series([10, 20, 30], index=['a', 'b', 'c']) s2 = pd.Series([1, 2, 4], index=['a', 'b', 'd']) s1 + s2 # a 11.0, b 22.0, c NaN, d NaN s1.add(s2, fill_value=0) # a 11, b 22, c 30, d 4
Methods: add(), sub(), mul(), div(). They accept fill_value, which the plain operators do not. Because NaN is a float, the result type becomes float.
head() and tail()
s.head(n) returns the first n values; s.tail(n) returns the last n. With no number, both use 5. They are handy to peek at a long Series without printing all of it.
s.head(2) # Asha 92, Ravi 75 s.tail(3) # Mohan 58, Zoya 81, Kabir 66 s.head() # first 5
Selection, indexing and slicing
Indexing (one value)
s['Zoya'] # by label โ 81 s.loc['Zoya'] # by label (clear way) s.iloc[3] # by position โ 81 s[['Asha', 'Kabir']] # a list of labels โ a smaller Series
Slicing (a range)
s[1:4] # positions 1, 2, 3 (end NOT included) s['Ravi':'Zoya'] # labels Ravi to Zoya (end IS included) s[::-1] # reverse order
Selection by condition (boolean indexing)
s[s > 70] # only values above 70 s[s == 58]
You can also change values: s['Mohan'] = 60 or s[1:3] = 0.
Try it: your own marks Series
Open Python (IDLE, Jupyter or any online Python). Make a dictionary of your 5 subjects and marks, then turn it into a Series. Before running each line, write your guess on paper: s.head(2), s[s > 80], s[1:3], s + 5. Run them and tick the guesses you got right. Then try the same in step 6 of the 3D.
Key formulas and definitions
- pd.Series(data, index=[...]) ยท data = list, ndarray, dict or scalar
- Dict โ keys become index ยท Scalar โ needs index, value repeats
- s1 + s2 matches labels; unmatched label โ NaN ยท s1.add(s2, fill_value=0)
- head(n) / tail(n): first / last n (default 5)
- Position slice s[a:b] excludes b ยท Label slice s['p':'q'] includes q
Worked examples
1. Create a Series of the days in the first three months using a dictionary, and print it.
d = {'Jan': 31, 'Feb': 28, 'Mar': 31}; s = pd.Series(d); print(s) โ Jan 31, Feb 28, Mar 31, dtype: int64. The keys became the index.
2. Create a Series of four 100s with index 'w', 'x', 'y', 'z'.
s = pd.Series(100, index=['w', 'x', 'y', 'z']). The scalar 100 is repeated for each label.
3. s = pd.Series([92, 75, 58, 81, 66], index=['Asha','Ravi','Mohan','Zoya','Kabir']). What do s[1:4] and s['Ravi':'Zoya'] give?
Both give Ravi 75, Mohan 58, Zoya 81. Position slice 1:4 takes positions 1, 2, 3 (4 left out). Label slice includes the end label Zoya.
4. For the same s, write code to show only students with marks above 70, and count them.
s[s > 70] โ Asha 92, Ravi 75, Zoya 81. Count: s[s > 70].size or len(s[s > 70]) โ 3.
5. s1 = pd.Series([10, 20, 30], index=['a','b','c']); s2 = pd.Series([1, 2, 4], index=['a','b','d']). Find s1 + s2 and s1.add(s2, fill_value=0).
s1 + s2 โ a 11.0, b 22.0, c NaN, d NaN (c and d have no partner). With fill_value=0: a 11, b 22, c 30, d 4.
6. Using the marks Series s, give 5 grace marks to everyone below 70 in one line.
s[s < 70] = s[s < 70] + 5 โ Mohan 58 โ 63, Kabir 66 โ 71; others unchanged.
Common mistakes
- Thinking s1 + s2 adds by position; pandas adds by index label, and unmatched labels give NaN.
- Forgetting that a position slice s[1:4] leaves out position 4, while a label slice s['Ravi':'Zoya'] includes Zoya.
- Making a Series from a scalar without an index and expecting many values.
- Writing s.head[3] with square brackets; head is a method, so write s.head(3).