Exit Chapter 02 Β· πŸ“Š Data Handling Using Pandas – II β€” Quiz

Playing as a guest? to earn XP and track your progress.

Question 1 of 50 0 correct so far

πŸ“Š Data Handling Using Pandas – II Quiz

Q1.

Which function is used to iterate over a DataFrame row by row?

Explanation

The iterrows() function iterates over the DataFrame horizontally, returning the index and the row as a Series.

Q2.

What does the iteritems() function iterate over?

Explanation

iteritems() iterates over the DataFrame vertically, returning the column name and the column data as a Series.

Q3.

When performing binary operations between two DataFrames, how does Pandas align the data?

Explanation

Pandas aligns data based on both row indices and column labels before performing the operation.

Q4.

What is the result of adding two DataFrames if a label exists in one but not the other?

Explanation

If a label does not match in both DataFrames during a binary operation, the result for that location is NaN (Not a Number).

Q5.

Which function is used to calculate the arithmetic mean of a DataFrame?

Explanation

The mean() function calculates the arithmetic average of the values.

Q6.

By default, which axis does df.sum() operate on?

Explanation

By default, descriptive statistics functions like sum() operate on axis=0, which means they calculate values down the column.

Q7.

Which function returns the middle value of a distribution?

Explanation

median() sorts the data and finds the middle value, separating the higher half from the lower half.

Q8.

What does the mode() function return?

Explanation

mode() identifies and returns the value(s) that appear most frequently in the dataset.

Q9.

If mode() finds a tie (multiple values with same highest frequency), what is the return type?

Explanation

If there are multiple modes, Pandas returns a Series or DataFrame containing all of them, not a single scalar value.

Q10.

Which function counts the number of non-null observations?

Explanation

count() returns the number of non-NaN (valid) values. size returns the total number of elements including NaNs.

Q11.

What argument is used in df.sum() to calculate the sum of each row?

Explanation

Setting axis=1 performs the operation horizontally across columns, resulting in a sum for each row.

Q12.

What does df.quantile(0.5) is equivalent to?

Explanation

The 0.5 quantile represents the 50th percentile, which is mathematically the median.

Q13.

Which function computes the standard deviation?

Explanation

std() computes the standard deviation, which measures the amount of variation or dispersion of a set of values.

Q14.

What does the describe() function provide?

Explanation

describe() generates descriptive statistics that summarize the central tendency, dispersion, and shape of a dataset’s distribution.

Q15.

Which of the following is NOT included in the output of describe() for numeric data?

Explanation

describe() includes count, mean, std, min, 25%, 50%, 75%, and max. It does not explicitly list variance (var).

Q16.

Which function is used to print a concise summary of a DataFrame, including index dtype and column dtypes?

Explanation

info() prints a concise summary of a DataFrame, including the index dtype, column dtypes, non-null values, and memory usage.

Q17.

How do you view the top 5 rows of a DataFrame named df?

Explanation

head() displays the first n rows. If n is not specified, it defaults to 5.

Q18.

What will df.tail(3) return?

Explanation

tail(n) returns the last n rows of the DataFrame.

Q19.

Which function is used to calculate the cumulative sum of the DataFrame?

Explanation

cumsum() returns the cumulative sum over a DataFrame or Series axis.

Q20.

How is missing data typically represented in Pandas for numeric columns?

Explanation

Pandas uses NaN (Not a Number) to represent missing or undefined numeric data.

Q21.

Which function checks for missing values and returns a boolean object?

Explanation

isnull() (and its alias isna()) returns a DataFrame of boolean values indicating whether each element is missing.

Q22.

Which function is used to remove rows containing missing values?

Explanation

dropna() is used to remove missing values. It can drop rows or columns based on the axis specified.

Q23.

In df.dropna(how='all'), what does the how='all' parameter specify?

Explanation

how='all' ensures that a row is dropped only if every element in that row is NaN.

Q24.

Which function is used to replace missing values with a specified value?

Explanation

fillna() is used to fill NA/NaN values using the specified method or value.

Q25.

What is the primary purpose of the groupby() function?

Explanation

groupby() involves a process of splitting the object, applying a function, and combining the results.

Q26.

When iterating with iterrows(), can you modify the DataFrame directly using the returned row variable?

Explanation

iterrows() returns a Series which is a copy of the row, not a view. Modifying this Series does not change the original DataFrame.

Q27.

Which method helps in finding the variance of the data?

Explanation

var() returns the variance of the values over the requested axis.

Q28.

If df has 10 rows and 2 columns, what is the shape of the object returned by df.count()?

Explanation

count() returns a Series with the count of non-NA values for each column. Since there are 2 columns, the shape is (2,).

Q29.

Which library must be imported to handle NaN explicitly (e.g., np.nan)?

Explanation

NaN is a special floating-point value defined in the NumPy library (np.nan).

Q30.

What is the default behavior of df.min() regarding NaN values?

Explanation

Most descriptive statistics methods in Pandas, like min(), exclude missing data (NaN) by default.

Q31.

In the split-apply-combine strategy, which function performs the 'split' step?

Explanation

groupby() is responsible for splitting the data into groups based on some keys.

Q32.

Which of the following creates a DataFrame df?

Explanation

pd.DataFrame() is the constructor used to create a DataFrame.

Q33.

What will df.fillna(method='ffill') do?

Explanation

ffill stands for 'forward fill', which propagates the last valid observation forward to next valid.

Q34.

If you want to fill missing values with the mean of the column, which code is correct?

Explanation

You must calculate the mean first using df.mean() and pass that result to fillna().

Q35.

What is the output of df.size?

Explanation

size returns an integer representing the number of elements in this object (Rows Γ— Columns).

Q36.

Which method allows applying a function along an axis of the DataFrame?

Explanation

apply() applies a function along an axis of the DataFrame.

Q37.

To perform element-wise addition of two DataFrames df1 and df2 filling missing values with 0, use:

Explanation

Using the method add() allows specifying parameters like fill_value, which + operator does not support.

Q38.

What is the data type of the index returned by iterrows()?

Explanation

iterrows() yields (index, Series), where index is the label of the row, matching the DataFrame's index type.

Q39.

Which function allows you to get the maximum value from a specific column 'A'?

Explanation

Selecting the column df['A'] returns a Series, and .max() calculates the maximum of that Series.

Q40.

What does df.isnull().sum() return?

Explanation

isnull() returns booleans (True/1 for NaN). Summing them counts the number of True values (missing values) per column.

Q41.

If you use df.groupby('Class')['Marks'].mean(), what is the index of the resulting Series?

Explanation

The column used for grouping ('Class') becomes the index of the resulting aggregated object.

Q42.

Which parameter in read_csv handles missing values while loading data?

Explanation

na_values is used to specify additional strings to recognize as NA/NaN.

Q43.

What is the default value of the axis parameter in df.drop()?

Explanation

By default, drop() removes rows (labels from index), so axis=0.

Q44.

Which function would you use to find the 25th percentile?

Explanation

quantile(q) calculates the value at the given quantile q (0 < q < 1). 0.25 is the 25th percentile.

Q45.

Can groupby() be used with multiple columns?

Explanation

Yes, you can group by multiple columns by passing them as a list, e.g., df.groupby(['School', 'Class']).

Q46.

What happens if you subtract a DataFrame from itself using df - df?

Explanation

Standard arithmetic rules apply: x - x = 0. However, NaN - NaN = NaN. So valid numbers become 0, NaNs stay NaN.

Q47.

Which attribute gives the axes labels of the DataFrame?

Explanation

df.axes returns a list representing the axes: the row index (axis 0) and column columns (axis 1).

Q48.

To sort the DataFrame by the index, use:

Explanation

sort_index() sorts the object by labels (along an axis).

Q49.

What is the return type of df['Column']?

Explanation

Selecting a single column from a DataFrame returns a pandas Series.

Q50.

Which symbol is used for matrix multiplication in newer Python versions (often supported by Pandas wrappers)?

Explanation

The @ operator is used for matrix multiplication. For element-wise multiplication, * is used.