Attribute types, data quality, and distance

Data Mining · Lecture 2 ·

A three-by-four right triangle gives Euclidean distance five.
Euclidean distance combines coordinate differences. Feature units and scaling determine how much each coordinate contributes.

The meaning of an attribute determines which comparisons and calculations are valid. A numeric-looking code may identify a category rather than measure a quantity. Data analysis begins by matching the representation to the underlying meaning.

Measurement scales #

Nominal attributes identify categories without an intrinsic order. Ordinal attributes have an order but not necessarily equal intervals. Interval attributes have meaningful differences but an arbitrary zero; ratio attributes also have a meaningful zero, allowing ratios to be interpreted.

Discrete attributes take separated values, such as counts. Continuous measurements can vary across a range even though recorded precision is finite. These distinctions answer different questions: an ordered rating can be discrete and ordinal at the same time.

Structure and quality #

Dimensionality counts features; sparsity describes how few entries are nonzero or present; resolution describes the level of detail. A larger dataset is not necessarily more informative if its observations are duplicated or inappropriate for the question.

Noise is unwanted variation, while an outlier is an observation far from a typical pattern. An outlier may be an error or the main phenomenon of interest. Missing values are not automatically zeros. A duplicate might be an accidental repetition or a legitimate second event. Cleaning should follow the meaning and provenance of the data.

Distance between observations #

Euclidean distance is the square root of the sum of squared coordinate differences. Manhattan distance is the sum of their absolute values. Between (1, 2) and (4, 6), the differences are 3 and 4, so the distances are 5 and 7 respectively.

A distance matrix records pairwise comparisons. Its diagonal is zero for an ordinary distance measure, and the matrix is symmetric when the measure is symmetric. Checking those properties helps catch calculation mistakes.

Scale and interpretation #

If one feature is measured in thousands and another in single digits, the larger-range feature can dominate a distance calculation. Scaling changes that balance, but it does not fix an irrelevant feature or an invalid measurement. Choose a comparison based on the data's meaning, then inspect whether its numeric behavior matches that intention.[1]

References

  1. ↑ scikit-learn: preprocessing .