Supervised learning and decision-tree rules

Data Mining · Lecture 8 ·

A decision tree sends values at or below 5 left and larger values right.
A decision tree applies a sequence of feature tests. Each leaf gives a prediction for the observations reaching it.

A decision tree predicts a category by following a sequence of feature tests. Internal nodes ask questions, branches represent their outcomes, and leaves supply predictions. The tree makes the relationship between measurements and a classification rule visible.[1]

Learning from labeled examples #

Supervised learning begins with examples containing both features and known labels. A tree divides those examples into groups so that the resulting groups have more consistent labels. An unsupervised task, by contrast, seeks structure without being given the desired label for each example.

A threshold question might ask whether a measurement is below a chosen value. Its two branches may lead directly to predictions or to further tests. To classify a new observation, follow the tests from the root until a leaf is reached.

A worked rule #

Consider a teaching example with two measurements. The root asks whether the first measurement is below 2. A yes branch predicts category A; the no branch asks whether the second measurement is below 5, predicting B or C. An observation (3, 4) takes the no branch and then the yes branch, producing B. This illustrates traversal, not a rule claimed to work for an actual dataset.

Choosing a split #

A useful split separates classes while leaving enough observations to support the result. Impurity measures summarize label mixture. A pure node contains one class; a mixed node contains several. Comparing the weighted impurity of child nodes helps evaluate whether a split improves separation.

Evaluation and limits #

A confusion matrix compares predicted categories with actual categories and shows which errors occur. A tree that fits the training rows perfectly can still fail on new examples if it has learned accidental details. Separate evaluation data and appropriate limits on tree complexity help detect that problem.

References

  1. ↑ scikit-learn: decision trees .