Files
machine-learning/02_decision_tree/02_decision_tree.ipynb
T
2024-09-25 18:29:02 +08:00

45 KiB

决策树

目录

  • 决策树学习器

决策树学习器

概述

决策树

决策树是一个流程图,它使用决策树及其可能的后果进行分类。在树的每个非叶子节点上,输入的一个属性被测试,根据这个测试结果,选择通往子节点的相应分支。在叶子节点上,根据这个叶子节点的类别标签,对输入进行分类。从根到叶的路径代表分类规则,根据这些规则给叶节点分配类标签。

decision tree

决策树学习

决策树学习是指从有类标记的训练数据中构建决策树。数据预计是一个元组,其中元组的每个记录都是用于分类的属性。决策树是自上而下构建的,通过在每一步选择一个变量来最好地分割项目集。有不同的指标来衡量 "最佳分割"。这些指标通常衡量子集内目标变量的同质性。

信息增益

信息增益是基于信息理论中的熵的概念。熵的定义为:

H(p) = -\sum{p_i \log_2{p_i}}

信息增益是指父代的熵和子代的熵的加权和之间的差异。用于分割的特征是提供最大信息增益的特征。

伪代码

function DECISION-TREE-LEARNING(examples, attributes, parent_examples) returns a tree
 if examples 是空集 then return PLURALITY-VALUE(parent_examples)
 else if examples 分类结果都相同 then return 分类结果
 else if attributes 是空集 then return PLURALITY-VALUE(examples)
 else
   A ← argmaxa ∈ attributes IMPORTANCE(a, examples)
   tree ← 以特征 A 为根检测节点的决策树
   for each 特征 A 的取值 vk do
     exs ← { e : e ∈ examples and e.A = vk }
     subtree ← DECISION-TREE-LEARNING(exs, attributes − A, examples)
     将标签为 (A = vk) 的子树 subtree 添加为 tree 的分支
   return tree

实现

由我们的学习算法构建的树的节点,根据它们是内部节点还是叶节点,分别使用DecisionFork或DecisionLeaf来存储。

In [1]:
import sys
sys.path.insert(1, '../')
from utils.utils import *

from utils.dataset4learners import *
from DecisionTreeLearner_3 import *
In [2]:
psource(DecisionFork)

class DecisionFork:
    """
    A fork of a decision tree holds an attribute to test, and a dict
    of branches, one for each of the attribute's values.
    """

    def __init__(self, attr, attr_name=None, default_child=None, branches=None):
        """Initialize by saying what attribute this node tests."""
        self.attr = attr
        self.attr_name = attr_name or attr
        self.default_child = default_child
        self.branches = branches or {}

    def __call__(self, example):
        """Given an example, classify it using the attribute and the branches."""
        attr_val = example[self.attr]
        if attr_val in self.branches:
            return self.branches[attr_val](example)
        else:
            # return default class when attribute is unknown
            return self.default_child(example)

    def add(self, val, subtree):
        """Add a branch. If self.attr = val, go to the given subtree."""
        self.branches[val] = subtree

    def display(self, indent=0):
        name = self.attr_name
        print('Test', name)
        for (val, subtree) in self.branches.items():
            print(' ' * 4 * indent, name, '=', val, '==>', end=' ')
            subtree.display(indent + 1)

    def __repr__(self):
        return 'DecisionFork({0!r}, {1!r}, {2!r})'.format(self.attr, self.attr_name, self.branches)

DecisionFork持有属性,在该节点进行测试,以及一个分支的决定。分支存储了子节点,每个属性的值都有一个。以输入元组为参数,以函数形式调用这个类的对象,根据属性测试的结果返回分类路径中的下一个节点。

In [3]:
psource(DecisionLeaf)

class DecisionLeaf:
    """A leaf of a decision tree holds just a result."""

    def __init__(self, result):
        self.result = result

    def __call__(self, example):
        return self.result

    def display(self):
        print('RESULT =', self.result)

    def __repr__(self):
        return repr(self.result)

叶子节点在result中存储类别标签。所有输入图元的分类路径都在DecisionLeaf上结束,其result 属性决定了它们的类别。

In [4]:
psource(DecisionTreeLearner)

class DecisionTreeLearner:
    """DecisionTreeLearner: based on information gain"""

    def __init__(self, dataset):
        self.dataset = dataset
        self.tree = self.decision_tree_learning(dataset.examples, dataset.inputs)

    def decision_tree_learning(self, examples, attrs, parent_examples=()):
        raise NotImplementedError

    def plurality_value(self, examples):
        """
        Return the most popular target value for this set of examples.
        (If target is binary, this is the majority; otherwise plurality).
        """
        popular = argmax_random_tie(self.dataset.values[self.dataset.target],
                                    key=lambda v: self.count(self.dataset.target, v, examples))
        return DecisionLeaf(popular)

    def count(self, attr, val, examples):
        """Count the number of examples that have example[attr] = val."""
        return sum(e[attr] == val for e in examples)

    def all_same_class(self, examples):
        """Are all these examples in the same target class?"""
        class0 = examples[0][self.dataset.target]
        return all(e[self.dataset.target] == class0 for e in examples)

    def choose_attribute(self, attrs, examples):
        """Choose the attribute with the highest information gain."""
        return argmax_random_tie(attrs, key=lambda a: self.information_gain(a, examples))

    def information_gain(self, attr, examples):
        """Return the expected reduction in entropy from splitting by attr."""
        raise NotImplementedError

    def split_by(self, attr, examples):
        """Return a list of (val, examples) pairs for each val of attr."""
        return [(v, [e for e in examples if e[attr] == v]) for v in self.dataset.values[attr]]

    def predict(self, x):
        return self.tree(x)

    def __call__(self, x):
        return self.predict(x)

上面的实现使用信息增益作为衡量标准来选择测试哪一个属性进行拆分。该函数以递归的方式自上而下地构建树。根据输入,它做出四个选择中的一个。

  1. 如果当前步骤的输入没有训练数据,我们将返回在父步骤(上一级递归)中收到的输入数据的类别模式。
  2. 如果训练数据中的所有值都属于同一类别,它将返回一个DecisionLeaf,其类别标签是所有数据所属的类别。
  3. 如果数据没有可以测试的属性,我们就返回训练数据中具有最高复数值的类。
  4. 我们选择熵值最高的属性,并返回一个基于此属性的DecisionFork。每个分支递归地调用decision_tree_learning来构建子树。

实现要点

def information_content(values):
    """Number of bits to represent the probability distribution in values."""
    probabilities = values 的归一化数值
    return probabilities 代入信息熵公式

def information_gain(self, attr, examples):
    """Return the expected reduction in entropy from splitting by attr."""

    def I(examples):
        return information_content([self.count(self.dataset.target, v, examples)
                                    for v in self.dataset.values[self.dataset.target]])

    n = 样本数
    remainder = 剩余信息熵值
    return 信息增益

例子

现在我们将使用决策树学习器对一个有数值的样本进行分类:5.1, 3.0, 1.1, 0.1.

In [5]:
iris = DataSet(name="iris")
DTL = DecisionTreeLearner(iris)
print(DTL([5.1, 3.0, 1.1, 0.1]))
#print(DTL.predict([5.1, 3.0, 1.1, 0.1]))
setosa

正如预期的那样,决策树学习器将样本归类为 "setosa"。

In [6]:
assert DTL.predict([5, 3, 1, 0.1]) == 'setosa'
assert DTL.predict([6, 5, 3, 1.5]) == 'versicolor'
assert DTL.predict([7.5, 4, 6, 2]) == 'virginica'