Data cleaning, feature engineering, and classification

Data Mining · Lecture 7 ·

A box plot labels the first quartile, median, third quartile, and interquartile range.
The interquartile range is Q3 minus Q1. It summarizes the spread of the middle half of the observations.

A reproducible cleaning workflow transforms a raw table into a documented modeling dataset. The central decisions concern which observations to retain, how to represent missingness, and which columns can validly predict the target.

Targeted cleanup #

Begin with column names, types, summary statistics, and missing-value counts. Drop a row when a missing required field makes that row unusable for the task, rather than automatically dropping every row with any missing field. Median imputation may be appropriate for some numeric columns; an explicit unknown category may be more interpretable for a categorical field.[1]

Filtering and duplicate removal change the dataset's composition. Record the criteria and compare row counts before and after each step so that an unexpected loss of data is visible.

Engineering a numeric duration #

A duration stored as text combines a value and a unit. Extracting digits alone can make unlike quantities appear comparable: minutes and seasons are not the same measurement. Restrict to a consistent unit or retain the unit as part of the representation before converting to a numeric type.

Interquartile range #

The interquartile range is Q3 - Q1, the distance between the upper and lower quartiles. A common screening rule flags observations below Q1 - 1.5 × IQR or above Q3 + 1.5 × IQR. These fences identify unusual values under that rule; they do not prove that a value is erroneous. A long-duration observation might be legitimate and important.

Features and labels #

Supervised classification uses input features, often written X, and known category labels, often written y. A binary classifier distinguishes two categories; multiclass classification distinguishes more. Training examples teach a rule, while separate evaluation examples test its performance on data not used to construct it.

A target or a direct consequence of the target must not accidentally appear among predictors when it would be unavailable at prediction time. Such leakage can produce impressive but unusable accuracy.

Visualization and interpretation #

Plots reveal distributions, category imbalance, and suspicious relationships before and after cleaning. A figure complements numeric summaries rather than replacing them. The final model's reliability depends on the whole preparation process, including decisions that occurred before the first model-fitting call.

References

  1. ↑ pandas user guide: Working with missing data .