Why files? Text, binary and CSV files
Data in variables lives in RAM and is lost when the program ends. A file stores data on disk permanently.
- Text file (.txt): human-readable characters. Each line ends with an EOL (end of line) character
\n. You can open it in Notepad. - Binary file (.dat, .bin): raw bytes in the same form as in memory. Not readable in Notepad. Fast, and can store Python objects (via pickle). No EOL translation.
- CSV file (.csv, comma-separated values): a text file where each line is one row of a table and values are separated by commas, like
1,Asha,92. Opens in spreadsheets.
Absolute and relative paths
A path tells Python where the file is.
- Absolute path: full address from the root of the drive, like
C:\\School\\CS\\marks.txtor/home/asha/cs/marks.txt. - Relative path: address from the current working directory (the folder the program runs in).
marks.txtmeans "in this folder";data/marks.txtmeans "in the data sub-folder";../marks.txtmeans "one folder up".
In Windows paths, write r'C:\School\marks.txt' (raw string) or use double backslashes, because \n or \t inside a path would be read as special characters.
Opening and closing a file: open modes and with
f = open('marks.txt', 'r') returns a file object (file handle). Always f.close() when done; closing saves any data still waiting in the buffer.
| Mode | Meaning | If file missing | Cursor starts |
|---|---|---|---|
| r | read only (default) | FileNotFoundError | start |
| w | write only, wipes old data | creates it | start |
| a | append (add at end) | creates it | end |
| r+ | read and write | error | start |
| w+ | write and read, wipes | creates it | start |
| a+ | append and read | creates it | end |
Add b for binary files: rb, wb, ab, rb+, wb+, ab+.
The with statement
with open('marks.txt', 'r') as f:
data = f.read()
# the file is closed here automaticallywith closes the file by itself when the block ends, even if an exception occurs.
Writing to a text file: write and writelines
f.write(s)writes one string and returns the number of characters written. It does not add a newline; add'\n'yourself.f.writelines(L)writes every string of a list (or other sequence) one after another, also without adding newlines.
with open('poem.txt', 'w') as f:
f.write('Roses are red\n')
f.writelines(['Sky is blue\n', 'Python is fun\n'])To add without wiping, open in 'a' mode.
Reading a text file: read, readline, readlines
f.read()reads the whole rest of the file as one string.f.read(n)reads at most n characters.f.readline()reads one line, including its\n. At the end of the file it returns''.f.readlines()reads all lines into a list of strings.- Loop line by line:
for line in f:
# count lines that start with 'T'
c = 0
with open('poem.txt') as f:
for line in f:
if line.startswith('T'):
c += 1
print(c)Use line.strip() to remove the trailing newline, and line.split() to break a line into words.
seek() and tell(): moving the file pointer
The file pointer is the position (in bytes from the start) where the next read or write happens.
f.tell()returns the current position.f.seek(offset, whence)moves it.whence0 = from start (default), 1 = from current position, 2 = from end. In text mode onlywhence = 0is safe with non-zero offsets; in binary mode all three work.
with open('poem.txt') as f:
print(f.read(5)) # 'Roses'
print(f.tell()) # 5
f.seek(0)
print(f.read(3)) # 'Ros'
Binary files with pickle: dump and load
Pickling (serialisation) turns a Python object into a stream of bytes. Unpickling turns bytes back into the object.
import pickle
rec = {'roll': 1, 'name': 'Asha', 'marks': 92}
with open('stu.dat', 'wb') as f:
pickle.dump(rec, f)
with open('stu.dat', 'rb') as f:
r = pickle.load(f)
print(r['name']) # AshaEach dump writes one object. To read all objects, call load in a loop until EOFError:
with open('stu.dat', 'rb') as f:
try:
while True:
print(pickle.load(f))
except EOFError:
pass
Search, append and update in a binary file
Append
Open in 'ab' and pickle.dump the new record. Old records stay.
Search
Open in 'rb', load records one by one, and check a field: if r['roll'] == 5: print(r); found = True. After the loop, print 'Not found' if found is still False.
Update
Method 1 (simple): load all records into a list, change the matching one, then open the file in 'wb' and dump all again.
recs = []
with open('stu.dat', 'rb') as f:
try:
while True:
recs.append(pickle.load(f))
except EOFError:
pass
for r in recs:
if r['roll'] == 2:
r['marks'] = 80
with open('stu.dat', 'wb') as f:
for r in recs:
pickle.dump(r, f)Method 2 (in place): open in 'rb+', note pos = f.tell() before each load, and when the record matches, f.seek(pos) and dump the changed record (works when the new record has the same size in bytes).
CSV files: csv.writer and csv.reader
import csv
with open('class.csv', 'w', newline='') as f:
w = csv.writer(f)
w.writerow(['Roll', 'Name', 'Marks'])
w.writerows([[1, 'Asha', 92], [2, 'Ravi', 75]])
with open('class.csv', 'r') as f:
for row in csv.reader(f):
print(row) # ['1', 'Asha', '92'] …csv.writer(f)makes a writer object.writerow(list)writes one row;writerows(list_of_lists)writes many.newline=''stops blank lines between rows on Windows.csv.reader(f)gives each row as a list of strings; convert numbers withint().delimiter=';'can change the separator.
Try it: a mini diary
Write a program that asks for one line and appends it to diary.txt using mode 'a'. Run it three times. Now open the file in mode 'w' once and write one line. Open diary.txt in Notepad after each run. You will see with your own eyes that 'w' wiped everything. Then use the free-play step in the 3D to predict what seek and read(n) return.
Key formulas and definitions
- open(path, mode) · modes: r, w, a, r+, w+, a+ (+ b for binary)
- with open(...) as f: → file closes by itself
- read(n), readline(), readlines() · write(s), writelines(list)
- tell() → position · seek(offset, whence) whence 0 start, 1 current, 2 end
- pickle.dump(obj, f) / obj = pickle.load(f) · loop until EOFError
- csv.writer(f).writerow(row) / writerows(rows) · csv.reader(f) → lists of strings
Worked examples
1. Which mode would you use: (a) add new rows to a log without losing old ones (b) make a new report file every time (c) read a saved binary record?
(a) 'a' (append; cursor at end). (b) 'w' (creates or wipes). (c) 'rb' (read binary).
2. poem.txt holds 'Roses are red\nSky is blue\n'. What do f.readline(), then f.read(3), then f.tell() give?
readline() → 'Roses are red\n' (14 characters: 13 letters/spaces + newline). read(3) → 'Sky'. tell() → 14 + 3 = 17 (on a system where newline is 1 byte).
3. Write a function that counts the words in story.txt.
def count_words(): with open('story.txt') as f: return len(f.read().split()) print(count_words())
4. Write a function that displays lines of notes.txt that are shorter than 20 characters.
def short_lines(): with open('notes.txt') as f: for line in f: if len(line.rstrip('\n')) < 20: print(line, end='')
5. A binary file book.dat has records [book_no, title, price]. Write a function to add one record.
import pickle def add_book(): rec = [int(input('No: ')), input('Title: '), float(input('Price: '))] with open('book.dat', 'ab') as f: pickle.dump(rec, f)
6. Write a function that shows all books costing more than 300 from book.dat and says if none were found.
def costly(): found = False with open('book.dat', 'rb') as f: try: while True: r = pickle.load(f) if r[2] > 300: print(r); found = True except EOFError: pass if not found: print('No such book')
7. Update the price of book number 3 to 450 in book.dat.
Load all records into a list (loop until EOFError), set r[2] = 450 where r[0] == 3, then reopen with 'wb' and dump every record again. Mode 'wb' is needed because we rewrite the whole file.
8. Write two functions: add rows [id, name, phone] to contacts.csv, and print the names of all contacts.
import csv def add(rows): with open('contacts.csv', 'a', newline='') as f: csv.writer(f).writerows(rows) def names(): with open('contacts.csv') as f: for row in csv.reader(f): print(row[1])
Common mistakes
- Opening an existing file in 'w' mode to add data; it wipes everything. Use 'a'.
- Forgetting that write() and writelines() do not add '\n' by themselves.
- Using pickle with a text-mode file; pickle needs 'wb', 'rb' or 'ab'.
- Expecting csv.reader to return numbers; it returns strings, so convert with int() or float().