Playing as a guest? to earn XP and track your progress.
π Data Handling Using Pandas β II Quiz
π Data Handling Using Pandas β II
Which function is used to iterate over a DataFrame row by row?
iterrows()
iteritems()
iterrow()
iter()
The iterrows() function iterates over the DataFrame horizontally, returning the index and the row as a Series.
What does the iteritems() function iterate over?
Rows
Columns
Elements
Indices
iteritems() iterates over the DataFrame vertically, returning the column name and the column data as a Series.
When performing binary operations between two DataFrames, how does Pandas align the data?
By row index only
By column label only
By both row index and column label
By position/integer index
Pandas aligns data based on both row indices and column labels before performing the operation.
What is the result of adding two DataFrames if a label exists in one but not the other?
0
Error
NaN
The value from the existing DataFrame
If a label does not match in both DataFrames during a binary operation, the result for that location is NaN (Not a Number).
Which function is used to calculate the arithmetic mean of a DataFrame?
average()
mean()
avg()
mode()
The mean() function calculates the arithmetic average of the values.
By default, which axis does df.sum() operate on?
axis=0 (column-wise)
axis=1 (row-wise)
axis=2
It requires an argument
By default, descriptive statistics functions like sum() operate on axis=0, which means they calculate values down the column.
Which function returns the middle value of a distribution?
mean()
median()
mode()
mid()
median() sorts the data and finds the middle value, separating the higher half from the lower half.
What does the mode() function return?
The average value
The most frequently occurring value
The middle value
The standard deviation
mode() identifies and returns the value(s) that appear most frequently in the dataset.
If mode() finds a tie (multiple values with same highest frequency), what is the return type?
Integer
Float
DataFrame/Series
Error
If there are multiple modes, Pandas returns a Series or DataFrame containing all of them, not a single scalar value.
Which function counts the number of non-null observations?
size()
count()
len()
shape()
count() returns the number of non-NaN (valid) values. size returns the total number of elements including NaNs.
What argument is used in df.sum() to calculate the sum of each row?
axis=0
axis=1
row=True
index=1
Setting axis=1 performs the operation horizontally across columns, resulting in a sum for each row.
What does df.quantile(0.5) is equivalent to?
df.mean()
df.mode()
df.median()
df.std()
The 0.5 quantile represents the 50th percentile, which is mathematically the median.
Which function computes the standard deviation?
var()
dev()
std()
stand()
std() computes the standard deviation, which measures the amount of variation or dispersion of a set of values.
What does the describe() function provide?
Only the mean
A summary of central tendency, dispersion, and shape
Only the data types
The first 5 rows
describe() generates descriptive statistics that summarize the central tendency, dispersion, and shape of a datasetβs distribution.
Which of the following is NOT included in the output of describe() for numeric data?
mean
std
variance
max
describe() includes count, mean, std, min, 25%, 50%, 75%, and max. It does not explicitly list variance (var).
Which function is used to print a concise summary of a DataFrame, including index dtype and column dtypes?
describe()
summary()
info()
details()
info() prints a concise summary of a DataFrame, including the index dtype, column dtypes, non-null values, and memory usage.
How do you view the top 5 rows of a DataFrame named df?
df.top()
df.head()
df.first(5)
df.rows(5)
head() displays the first n rows. If n is not specified, it defaults to 5.
What will df.tail(3) return?
The first 3 rows
The last 3 rows
The 3rd row
All rows except the last 3
tail(n) returns the last n rows of the DataFrame.
Which function is used to calculate the cumulative sum of the DataFrame?
sum(cumulative=True)
cumsum()
totalsum()
accumulate()
cumsum() returns the cumulative sum over a DataFrame or Series axis.
How is missing data typically represented in Pandas for numeric columns?
Null
None
NaN
Empty
Pandas uses NaN (Not a Number) to represent missing or undefined numeric data.
Which function checks for missing values and returns a boolean object?
checknull()
isnull()
isnan()
missing()
isnull() (and its alias isna()) returns a DataFrame of boolean values indicating whether each element is missing.
Which function is used to remove rows containing missing values?
delna()
removena()
dropna()
cleanna()
dropna() is used to remove missing values. It can drop rows or columns based on the axis specified.
In df.dropna(how='all'), what does the how='all' parameter specify?
Drop row if any value is NaN
Drop row only if all values are NaN
Drop all rows
Drop all columns
how='all' ensures that a row is dropped only if every element in that row is NaN.
Which function is used to replace missing values with a specified value?
replace()
fillna()
putna()
insert()
fillna() is used to fill NA/NaN values using the specified method or value.
What is the primary purpose of the groupby() function?
To sort data
To split data into groups based on criteria
To merge two DataFrames
To plot graphs
groupby() involves a process of splitting the object, applying a function, and combining the results.
When iterating with iterrows(), can you modify the DataFrame directly using the returned row variable?
Yes, it modifies the original DataFrame
No, it returns a copy
Yes, but only for numeric columns
No, it raises an error
iterrows() returns a Series which is a copy of the row, not a view. Modifying this Series does not change the original DataFrame.
Which method helps in finding the variance of the data?
std()
var()
cov()
variance()
var() returns the variance of the values over the requested axis.
If df has 10 rows and 2 columns, what is the shape of the object returned by df.count()?
(10, 2)
(10,)
(2,)
Scalar value
count() returns a Series with the count of non-NA values for each column. Since there are 2 columns, the shape is (2,).
Which library must be imported to handle NaN explicitly (e.g., np.nan)?
pandas
math
numpy
sys
NaN is a special floating-point value defined in the NumPy library (np.nan).
What is the default behavior of df.min() regarding NaN values?
It includes them as 0
It skips them
It returns NaN if any value is NaN
It raises an error
Most descriptive statistics methods in Pandas, like min(), exclude missing data (NaN) by default.
In the split-apply-combine strategy, which function performs the 'split' step?
apply()
transform()
groupby()
agg()
groupby() is responsible for splitting the data into groups based on some keys.
Which of the following creates a DataFrame df?
pd.DataFrame()
pd.CreateDataFrame()
pd.df()
pd.Table()
pd.DataFrame() is the constructor used to create a DataFrame.
What will df.fillna(method='ffill') do?
Fill with 0
Fill with the next value
Propagate the last valid observation forward
Fill with the mean
ffill stands for 'forward fill', which propagates the last valid observation forward to next valid.
If you want to fill missing values with the mean of the column, which code is correct?
df.fillna('mean')
df.fillna(df.mean())
df.replace(NaN, mean)
df.dropna(mean)
You must calculate the mean first using df.mean() and pass that result to fillna().
What is the output of df.size?
Number of rows
Number of columns
Total number of elements (rows * columns)
Memory usage
size returns an integer representing the number of elements in this object (Rows Γ Columns).
Which method allows applying a function along an axis of the DataFrame?
map()
apply()
iter()
call()
apply() applies a function along an axis of the DataFrame.
To perform element-wise addition of two DataFrames df1 and df2 filling missing values with 0, use:
df1 + df2
df1.add(df2, fill_value=0)
df1.sum(df2)
df1.append(df2)
Using the method add() allows specifying parameters like fill_value, which + operator does not support.
What is the data type of the index returned by iterrows()?
Integer
String
Same as the DataFrame's index
Tuple
iterrows() yields (index, Series), where index is the label of the row, matching the DataFrame's index type.
Which function allows you to get the maximum value from a specific column 'A'?
df.max('A')
df['A'].max()
df.maximum('A')
max(df['A'])
Selecting the column df['A'] returns a Series, and .max() calculates the maximum of that Series.
What does df.isnull().sum() return?
Total number of values
Total number of missing values in each column
Total sum of all values
Boolean table
isnull() returns booleans (True/1 for NaN). Summing them counts the number of True values (missing values) per column.
If you use df.groupby('Class')['Marks'].mean(), what is the index of the resulting Series?
0, 1, 2...
Marks
Class
Mean
The column used for grouping ('Class') becomes the index of the resulting aggregated object.
Which parameter in read_csv handles missing values while loading data?
na_values
missing
nan_rep
drop_na
na_values is used to specify additional strings to recognize as NA/NaN.
What is the default value of the axis parameter in df.drop()?
0 (Rows)
1 (Columns)
None
All
By default, drop() removes rows (labels from index), so axis=0.
Which function would you use to find the 25th percentile?
quantile(0.25)
percentile(25)
quarter(1)
mean(0.25)
quantile(q) calculates the value at the given quantile q (0 < q < 1). 0.25 is the 25th percentile.
Can groupby() be used with multiple columns?
No, only one column
Yes, by passing a list of column names
Yes, but only with numeric columns
No, it causes an error
Yes, you can group by multiple columns by passing them as a list, e.g., df.groupby(['School', 'Class']).
What happens if you subtract a DataFrame from itself using df - df?
All values become 0
All values become NaN
All values become 0, but NaNs remain NaN
Error
Standard arithmetic rules apply: x - x = 0. However, NaN - NaN = NaN. So valid numbers become 0, NaNs stay NaN.
Which attribute gives the axes labels of the DataFrame?
df.labels
df.axes
df.indices
df.headers
df.axes returns a list representing the axes: the row index (axis 0) and column columns (axis 1).
To sort the DataFrame by the index, use:
sort_values()
sort_index()
order_by()
arrange()
sort_index() sorts the object by labels (along an axis).
What is the return type of df['Column']?
DataFrame
Series
List
Array
Selecting a single column from a DataFrame returns a pandas Series.
Which symbol is used for matrix multiplication in newer Python versions (often supported by Pandas wrappers)?
*
**
@
&
The @ operator is used for matrix multiplication. For element-wise multiplication, * is used.