Preprocessing translates observations into a form that an analysis method can use. Encoding represents categories, scaling changes numeric ranges, and imputation supplies a chosen treatment for missing values. Python provides the control flow and data structures used to express these transformations.
Encoding categories #
An ordinal encoding assigns ordered values to ordered categories. Applying the same technique to unordered categories can imply a distance or rank that does not exist. One-hot encoding instead creates indicator columns: each category is represented by whether its indicator is present. This makes the category distinction explicit without assuming that category three is greater than category two.[1]
Scaling formulas #
Mean normalization subtracts the mean and divides by the observed range. Standardization subtracts the mean and divides by the standard deviation. These are different transformations:
mean-normalized x = (x - mean) / (maximum - minimum)
standardized x = (x - mean) / standard deviation
A feature with no variation has a zero denominator and needs a deliberate treatment. Scaling parameters should be learned from training data and then applied consistently to evaluation data; otherwise information from the evaluation set can leak into training.
Missingness and new features #
Imputation may use a mean, median, category, or model-based estimate. The choice depends on the attribute and the missingness process. Feature engineering creates a more useful representation, such as converting a duration string into a numeric amount. Preserve units and distinguish an unknown value from a genuine zero.
Python execution #
Indentation defines blocks. Variables are names bound to objects; objects have types, and operations must be appropriate for them. Converting a numeric string to a number requires an explicit conversion such as float(value). Importing a module provides a namespace for its functions and types.
values = [float(item) for item in ["2.5", "3.5"]]
mean = sum(values) / len(values)This small example produces 3.0. An empty list would require separate handling. Functions package transformations, while notebook text cells record assumptions and interpretation. Rerunning an analysis from the beginning helps reveal hidden dependence on earlier interactive state.