NumPy arrays and pandas tables

Data Mining · Lecture 6 ·

A table contains labeled rows and columns for city and temperature.
A pandas DataFrame organizes observations by row and variables by column, with explicit row and column labels.

NumPy represents numerical data in arrays, while pandas organizes labeled one-dimensional and tabular data. Their value comes from applying operations consistently across collections while retaining enough structure to interpret the result.

Array structure #

A NumPy array has a shape, a number of dimensions, a size, and a data type. Shape describes the length of each axis; size counts all elements. Elementwise arithmetic applies an operation to corresponding positions, unlike ordinary list concatenation or repetition.[1]

import numpy as np
values = np.array([2.0, 4.0, 6.0])
doubled = values * 2
average = values.mean()

The output values are 4, 8, and 12, while the mean of the original array is 4. Aggregation along an axis requires knowing what that axis represents. A vectorized expression can replace a manual element loop, but valid shapes and units remain necessary.

Series and DataFrames #

A pandas Series combines values with an index. A DataFrame combines labeled columns and a row index. Labels are meaningful: selecting a column by name and selecting a position are different operations. Aligning data by labels can be useful, but unexpected labels can also explain surprising missing results.

Inspect before cleaning #

Inspect column names, types, sample rows, summary statistics, and missing-value counts. A column that looks numeric may contain strings, and an apparent blank may use a special missing-value marker. Understanding the schema should precede any wholesale dropping or filling of rows.

Missing values, duplicates, and selection #

A cleanup method often returns a new object, so retain the returned result when that is the intended change. Duplicate removal requires specifying what makes two observations duplicates. A repeated value in one column does not necessarily mean a repeated record.

Boolean selection chooses rows for which a condition is true. Conditions must use the appropriate elementwise operators and grouping. After filtering, check the number of retained rows and the meaning of the index rather than assuming it has been reset to consecutive positions.

These tools simplify collection operations; they do not decide the substantive meaning of a missing observation, a duplicate, or an outlier.

References

  1. ↑ NumPy: array fundamentals .