Python CSV & Pandas: Student Guide to DataFrames
A student-friendly Grade 10 guide to Python Pandas, CSV files, DataFrames, filtering, and saving data, with examples and practice questions.
Python CSV Files & Pandas
Learn how to store tabular data, read it into Python, find the rows you need, collect new records, and save your results. Work through the examples, then test yourself with the practice questions.
1. What is a CSV file?
CSV means Comma-Separated Values. It is a simple text file used to store information in rows and columns. Each line usually represents one record, and commas separate the values.
For example, a file named students.csv might contain:
Roll,Name,Age,Class,City,Marks
101,ram,16,10,Kathmandu,85
102,Sita,15,10,Pokhara,92
103,HARI,17,10,Lalitpur,68
104,ram,15,10,Bhaktapur,74
105,Gita,16,10,Kathmandu,90
The first line contains the column headings. The lines below it contain student records. CSV files are useful for moving data between programs and keeping simple records.
2. What is a Pandas DataFrame?
Pandas is a Python library for working with structured data. A DataFrame is a table made of rows and columns—similar to a spreadsheet.
| Roll | Name | City | Marks |
|---|---|---|---|
| 101 | ram | Kathmandu | 85 |
| 102 | Sita | Pokhara | 92 |
| 103 | HARI | Lalitpur | 68 |
Before using Pandas, import it. The common short name is pd:
import pandas as pd
Here, import makes the library available, and as pd gives it a shorter name so we can write pd.read_csv() instead of typing the full library name.
3. Read a CSV file into Python
Use pd.read_csv() to load a CSV file into a DataFrame.
import pandas as pd
df = pd.read_csv("students.csv")
print(df.head())
pd.read_csv("students.csv")Opens and reads the CSV file. Pandas normally treats the first row as the column headings.
df = ...Stores the resulting table in a variable named
df. You can choose another variable name.print(df.head())Displays the first five rows by default, helping you check whether the file loaded correctly.
students.csv in the location from which your Python program expects to read it, or provide the correct file path.4. Filter rows to find information
Filtering means selecting only the rows that match a condition. In Pandas, a condition produces True or False values. Pandas keeps the rows where the condition is True.
Example A: Find marks greater than 80
high_scorers = df[df["Marks"] > 80]
print(high_scorers)
Read it from the inside out: df["Marks"] > 80 checks each student's marks. The outer df[ ... ] keeps only matching rows.
In the sample dataset, Sita (92), ram (85), and Gita (90) match this condition.
Example B: Match exact text
ktm_students = df[df["City"] == "Kathmandu"]
print(ktm_students)
The operator == asks whether a value is equal to the text "Kathmandu". It is different from =, which assigns a value to a variable. Text matching is case-sensitive, so "Kathmandu" and "kathmandu" are not identical strings.
Example C: Match a name regardless of letter case
ram_students = df[df["Name"].str.upper() == "RAM"]
print(ram_students)
.str.upper() converts text values in the Name column to uppercase for comparison. This allows names such as ram, Ram, and RAM to match "RAM".
Example D: Search for part of a word
result = df[df["Name"].str.contains("ra", case=False, na=False)]
print(result)
.str.contains("ra") searches within each name instead of requiring the whole name to match. case=False ignores letter case, and na=False treats missing values as non-matches.
5. Combine more than one condition
Sometimes one condition is not enough. Pandas lets you combine conditions using these operators:
| Operator | Meaning | Think of it as |
|---|---|---|
& |
AND: both conditions must be True | Condition 1 and condition 2 |
| |
OR: at least one condition must be True | Condition 1 or condition 2 |
Example: Kathmandu students with marks above 80
filtered_df = df[
(df["City"] == "Kathmandu") & (df["Marks"] > 80)
]
print(filtered_df)
The brackets around each condition are important. With &, a student must be from Kathmandu and have more than 80 marks. In our sample, ram (85) and Gita (90) match.
Example using OR
selected = df[
(df["City"] == "Pokhara") | (df["Marks"] > 90)
]
print(selected)
This keeps students who live in Pokhara or scored more than 90 (or both).
& and | to combine Pandas conditions, and put each complete condition inside parentheses. Do not replace them with Python's words and or or in this type of row-wise filter.6. Save filtered data to a new CSV file
After filtering, you may want to save the selected rows so you can use them later. Pandas uses .to_csv().
top_students = df[df["Marks"] >= 85]
top_students.to_csv("top_students.csv", index=False)
print("Filtered data saved!")
>= 85 means 85 or higher. The selected records are saved in top_students.csv.
What does index=False mean?
A DataFrame has a row index used by Pandas. Setting index=False tells Pandas not to add that index as an extra column in the exported CSV. If you use index=True (the default), the index is written to the file too.
| Code | What happens in the saved CSV? |
|---|---|
df.to_csv("file.csv", index=False) |
Does not write the DataFrame index. |
df.to_csv("file.csv", index=True) |
Writes the DataFrame index as an extra column. |
7. Create CSV data, take input, and combine records
You can also build a DataFrame directly in Python instead of reading an existing CSV file. A dictionary can store column names as keys and lists of values as the column data.
Type 1: Static data, without index
import pandas as pd
data = {
"Name": ["Alice", "Bob", "Charlie"],
"Score": [85, 90, 78]
}
df = pd.DataFrame(data)
df.to_csv("students_scores.csv", index=False)
print(df)
Static data means the values are already written in the program. pd.DataFrame(data) turns the dictionary into a table.
Type 2: Static data, with index
df.to_csv("students_scores_with_index.csv", index=True)
This writes the DataFrame index into the CSV file. Run it after creating df as shown above.
Type 3: Get a record from the user
import pandas as pd
user_name = input("Enter student name: ")
user_score = int(input("Enter student score: "))
data = {
"Name": [user_name],
"Score": [user_score]
}
df = pd.DataFrame(data)
df.to_csv("new_student.csv", index=False)
print(df)
input() collects text typed by the user. int() converts the score to an integer so it can be stored as a number.
Type 4: Combine existing records with a new record
import pandas as pd
static_data = {
"Name": ["Alice", "Bob"],
"Score": [85, 90]
}
df_static = pd.DataFrame(static_data)
user_name = input("Enter new student name: ")
user_score = int(input("Enter new student score: "))
df_user = pd.DataFrame({
"Name": [user_name],
"Score": [user_score]
})
df_combined = pd.concat([df_static, df_user], ignore_index=True)
df_combined.to_csv("all_students.csv", index=True)
print(df_combined)
pd.concat() joins DataFrames together. Here, it places the new student's row after Alice and Bob. ignore_index=True creates a fresh sequence of row indexes in the combined DataFrame. Notice that to_csv(..., index=True) then saves that index to the file.
static_data, or change index=True to index=False and compare the exported CSV files.8. Common mistakes to avoid
- Using
=instead of==: use==to compare a column value with a target. - Forgetting parentheses: write
(condition1) & (condition2)when combining filters. - Wrong column name: use the exact heading from the CSV, including spelling and capitalization, such as
"Marks". - Wrong file location: check that the CSV file exists where Python is looking for it.
- Unexpected index column: use
index=Falseif you do not want the DataFrame index included in the saved file. - Converting text to a number: use
int(input(...))when the input should be a whole number, such as a score.
9. Practice questions
Try answering before opening the solutions. The questions move from remembering terms to writing and understanding code.
Part A — Short-answer questions
1. What does CSV stand for?
Show answer
Comma-Separated Values.
2. What is a DataFrame in Pandas?
Show answer
A two-dimensional table-like structure containing rows and columns.
3. What is the purpose of pd.read_csv("students.csv")?
Show answer
It reads the CSV file and loads its data into a Pandas DataFrame.
4. Explain the difference between = and ==.
Show answer
= assigns a value; == compares two values for equality.
5. What is the difference between index=True and index=False in to_csv()?
Show answer
index=True writes the DataFrame index into the CSV; index=False leaves it out.
Part B — Read the code
6. Given marks 85, 92, 68, 74, and 90, which marks are selected by df[df["Marks"] > 80]?
Show answer
85, 92, and 90. The condition is strictly greater than 80.
7. Which rows match df[df["City"] == "Kathmandu"] in the sample dataset?
Show answer
Roll 101 (ram) and Roll 105 (Gita).
8. What does .str.upper() do in a filter?
Show answer
It converts text values in the selected column to uppercase for comparison.
9. In a combined filter, what does & mean? What does | mean?
Show answer
& means AND (both conditions must be true). | means OR (at least one condition must be true).
10. Why might a programmer use na=False in .str.contains()?
Show answer
It treats missing values as non-matches, helping the filter handle missing entries safely.
Part C — Write the code
11. Write code to read students.csv and display its first five rows.
Show answer
import pandas as pd
df = pd.read_csv("students.csv")
print(df.head())
12. Write a filter to find students whose marks are 85 or above.
Show answer
top_students = df[df["Marks"] >= 85]
13. Write a filter to find students from Pokhara.
Show answer
pokhara_students = df[df["City"] == "Pokhara"]
14. Write code to find Kathmandu students who scored more than 80.
Show answer
result = df[
(df["City"] == "Kathmandu") & (df["Marks"] > 80)
]
15. Create a DataFrame with names Mina and Arun and scores 88 and 91, then save it as scores.csv without the index.
Show answer
import pandas as pd
data = {
"Name": ["Mina", "Arun"],
"Score": [88, 91]
}
df = pd.DataFrame(data)
df.to_csv("scores.csv", index=False)
Part D — Practical challenge
16. Create a program that asks the user for a student's name and score, creates a DataFrame, and saves it to student_entry.csv without the index.
Show one possible solution
import pandas as pd
name = input("Enter student name: ")
score = int(input("Enter student score: "))
data = {"Name": [name], "Score": [score]}
df = pd.DataFrame(data)
df.to_csv("student_entry.csv", index=False)
print(df)
17. Start with two existing students, ask the user for a third student, combine the records with pd.concat(), and save the result. Which argument makes the combined DataFrame's index start fresh?
Show answer
Use pd.concat([df_static, df_user], ignore_index=True). The argument ignore_index=True creates a fresh index for the combined DataFrame.