Playing as a guest? to earn XP and track your progress.
π Data Handling Using Pandas β I Quiz
π Data Handling Using Pandas β I
Which library is primarily used for data manipulation and analysis in Python?
NumPy
Pandas
Matplotlib
Scikit-learn
Pandas is the core library designed for structured data manipulation and analysis.
What is the standard convention for importing the Pandas library?
import pandas as p
import pandas as pd
import pandas as pn
include pandas
The universal alias for importing Pandas is pd.
Which Pandas data structure represents a one-dimensional labeled array?
DataFrame
Series
Array
List
A Series is a 1-D labeled array capable of holding data of any type.
Which Pandas data structure represents a two-dimensional labeled data structure with columns of potentially different types?
Series
DataFrame
Matrix
Panel
A DataFrame is a 2-D data structure, similar to a table or spreadsheet.
What is the key difference between a Pandas Series and a standard Python list?
Series can hold multiple data types
Series has explicit index labels
List is faster
List is 2-dimensional
Unlike lists which use integer positions, Series supports custom index labels.
If you create a Series from a Python dictionary, what becomes the index?
The values of the dictionary
Integer sequence 0, 1, 2...
The keys of the dictionary
It must be specified manually
When creating a Series from a dict, the dictionary keys automatically become the index labels.
Which attribute returns the underlying data of a Series as a NumPy array?
s.data
s.array
s.values
s.list
s.values returns the data as a NumPy ndarray.
What does the shape attribute of a DataFrame return?
Total number of elements
(number of rows, number of columns)
(number of columns, number of rows)
Memory usage
df.shape returns a tuple representing the dimensionality: (rows, columns).
Which method is used to view the first 5 rows of a DataFrame by default?
df.view()
df.top()
df.head()
df.start()
df.head() returns the first n rows (default is 5).
How do you access a column named 'Marks' in a DataFrame df?
df('Marks')
df['Marks']
df.get('Marks')
df[Marks]
Columns are accessed using bracket notation with the column name as a string: df['Marks'].
What is the main difference between loc and iloc?
loc is for columns, iloc is for rows
loc is label-based, iloc is integer position-based
loc is faster than iloc
iloc includes the end index in slicing
loc selects data by label/name, whereas iloc selects by integer index/position.
When slicing using loc['A':'C'], is the end bound 'C' included?
Yes, it is inclusive
No, it is exclusive
Only if it is a number
Depends on the data type
Slicing with labels using loc is inclusive of both the start and stop bounds.
When slicing using iloc[0:3], which rows are selected?
Rows at indices 0, 1, 2, 3
Rows at indices 0, 1, 2
Rows at indices 1, 2, 3
Row 3 only
Slicing with iloc (integers) follows Python standard slicing: start is inclusive, end is exclusive. So, 0, 1, 2.
Which command is used to remove a column 'Grade' from DataFrame df?
df.remove('Grade')
df.drop('Grade', axis=0)
df.delete('Grade')
df.drop('Grade', axis=1)
drop is used to remove data. axis=1 specifies that we are dropping a column.
What function generates descriptive statistics (mean, std, min, max) for numeric columns?
df.stats()
df.info()
df.describe()
df.summary()
describe() computes summary statistics of the Series or DataFrame.
How would you filter a DataFrame df to show rows where 'Marks' > 80?
df['Marks' > 80]
df[df['Marks'] > 80]
df.filter('Marks' > 80)
df.query(Marks > 80)
Boolean indexing requires passing the condition df['Marks'] > 80 inside the brackets.
Which method sorts the DataFrame by the values of a specific column?
df.sort_index()
df.order_by()
df.sort_values()
df.arrange()
sort_values() sorts a DataFrame by its values (columns).
If you add two Series, how does Pandas handle indices that do not match?
It raises an error
It ignores mismatching indices
It fills them with NaN (Not a Number)
It stops the operation
Pandas aligns data by index. Indices present in one but not the other result in NaN.
What is the result of df.size?
Number of bytes used
Number of rows
Number of columns
Total number of elements (rows * columns)
size returns the total number of elements in the object.
Which attribute gives the index (row labels) of the DataFrame?
df.rows
df.labels
df.index
df.axes
df.index contains the labels for the rows.
To transpose a DataFrame (swap rows and columns), which attribute is used?
df.transpose
df.T
df.swap
df.flip
df.T is the accessor for the transpose of the DataFrame.
How do you check for missing values in a DataFrame?
df.missing()
df.isnull()
df.check_na()
df.empty()
isnull() (or isna()) returns a boolean mask indicating missing values.
Which code correctly creates a DataFrame from a dictionary of lists?
pd.DataFrame({'A': [1,2], 'B': [3,4]})
pd.DataFrame(['A': [1,2], 'B': [3,4]])
pd.DataFrame('A'=[1,2], 'B'=[3,4])
pd.DataFrame({[1,2], [3,4]})
The correct syntax passes a dictionary where keys are column names and values are lists of data.
What happens if you use pd.Series(5, index=['a', 'b', 'c'])?
Error: Data must be a list
Creates a Series with 5 in 'a' and NaN in others
Creates a Series with 5 repeated for every index
Creates a Series with index 0,1,2
This is called Scalar broadcasting. The value 5 is repeated for each label in the index.
Which argument in df.drop() specifies that we want to drop a row?
axis=1
axis=0
inline=True
row=True
axis=0 refers to the index (rows), while axis=1 refers to columns.
What allows Pandas to perform operations on entire arrays without loops?
Serialization
Vectorization
Normalization
Iteration
Vectorization allows applying operations to whole arrays/columns at once efficiently.
Which method provides a concise summary of a DataFrame including data types and non-null counts?
df.describe()
df.head()
df.info()
df.types()
df.info() prints information about the DataFrame including the index dtype, columns, non-null values and memory usage.
How can you rename columns in a DataFrame?
df.columns = {'Old':'New'}
df.rename(columns={'Old': 'New'})
df.change_name('Old', 'New')
df.replace('Old', 'New')
rename method with the columns parameter taking a dictionary is the standard way to rename specific columns.
If s is a Series, what does s.ndim return?
1
2
The length of the series
The type of data
A Series is strictly 1-dimensional, so ndim is always 1.
If df is a DataFrame, what does df.ndim return?
1
2
3
Variable
A DataFrame is strictly 2-dimensional, so ndim is always 2.
To select multiple columns 'A' and 'B', which syntax is correct?
df['A', 'B']
df[['A', 'B']]
df.cols('A', 'B')
df['A']['B']
You must pass a list of column names inside the brackets, resulting in double brackets: df[['A', 'B']].
What is the default index type if none is provided when creating a DataFrame?
RangeIndex (0, 1, 2...)
StringIndex
RandomIntegers
Empty
Pandas assigns a default RangeIndex starting from 0.
Which property indicates whether a Series is empty?
s.is_empty
s.empty
s.void
s.null
s.empty returns a boolean True if the Series/DataFrame contains no elements.
Which function allows applying a custom function to every element of a Series?
s.map()
s.apply()
s.run()
s.execute()
apply() invokes a function on values of the Series.
What does df.loc['A', 'B'] access?
Row 'A'
Column 'B'
The cell at Row 'A' and Column 'B'
Rows 'A' through 'B'
loc[row_label, column_label] accesses a specific scalar value.
How do you verify the data types of all columns in a DataFrame?
df.type
df.dtypes
df.format
df.class
df.dtypes returns the data type of each column in the DataFrame.
Which method allows you to fill missing values with a specific number?
df.fill()
df.fillna()
df.replace_nan()
df.fix()
fillna(value) is used to fill NA/NaN values using the specified method or value.
What implies that a pandas object is 'mutable'?
Its shape can be changed
Its values can be changed
It cannot be modified
It can only be read
Value mutability means you can modify the data contained in the structure.
Which operator is used for element-wise logical AND in Pandas filtering?
and
&
&&
with
In Pandas/NumPy boolean indexing, & is used for element-wise AND (unlike Python's and).
Which operator is used for element-wise logical OR in Pandas filtering?
or
|
||
plus
| is the bitwise OR operator used for element-wise logical OR in Pandas.
What error do you get if you try to access a column that does not exist?
ValueError
IndexError
KeyError
NameError
A KeyError is raised when a dictionary key (or column name) is not found.
Which method returns the last n rows of a DataFrame?
df.last(n)
df.bottom(n)
df.tail(n)
df.end(n)
tail(n) returns the last n rows.
Can a DataFrame column contain mixed data types (e.g., integers and strings)?
No, never
Yes, the dtype will be 'object'
Yes, the dtype will be 'mixed'
Only if specified explicitly
Yes, Pandas falls back to the 'object' dtype (generic Python objects) if types are mixed.
What is the output of len(df)?
Number of columns
Number of cells
Number of rows
Number of bytes
len() on a DataFrame returns the number of rows (length of the index).
To sort the index labels of a Series s, which method is used?
s.sort_values()
s.sort_index()
s.order()
s.rank()
sort_index() sorts the object by its index labels.
If you create pd.Series([10, 20]), what are the default index labels?
1, 2
0, 1
'0', '1'
No index
The default RangeIndex starts at 0. So indices are 0 and 1.
How do you add a new row to a DataFrame using loc?
df.add(index, data)
df.loc[new_index] = data
df.append(data)
df.insert(data)
Assigning data to a non-existent label using loc creates a new row: df.loc['new'] = ...
What does df.columns return?
A list of row labels
The data in the first column
An Index object containing column labels
The number of columns
columns attribute returns an Index object containing the column labels.
Which of the following is NOT a valid way to create a DataFrame?
From a list of dictionaries
From a dictionary of lists
From a NumPy array
From a simple integer
A single integer cannot be converted into a 2D DataFrame structure directly without context.
If s = pd.Series([1, 2, 3], index=['a', 'b', 'c']), what is s['b']?
1
2
3
Error
s['b'] accesses the value associated with label 'b', which is 2.