Skip to content

Validity indices

FCM needs the number of clusters \(c\) up front. Validity indices score a fitted partition so that you can compare different values of \(c\). All of them use the euclidean distance, work on the output of FCM (or a variant whose rows of u sum to one) and need \(c \geq 2\).

Function Better Measures
xie_beni lower compactness over separation of the centers
fukuyama_sugeno lower compactness minus spread of the centers around their mean
davies_bouldin lower membership weighted scatter over center distance
fuzzy_silhouette higher silhouette weighted by how clear each membership is

The equations and references are in the reference. fuzzy_silhouette builds the \(n \times n\) distance matrix, so it needs \(O(n^2)\) memory; the other three are linear in \(n\).

Usage

import numpy as np
from fcmeans import FCM, select_n_clusters, xie_beni

X = np.vstack([np.random.normal(loc=c, size=(50, 2)) for c in (0, 8, 16)])

fcm = FCM(n_clusters=3, random_state=42)
fcm.fit(X)
xie_beni(X, fcm.u, fcm.centers, fcm.m)

# fit for each c and recommend one
best, scores = select_n_clusters(
    X, range(2, 8), index="fuzzy_silhouette", random_state=42
)

Extra keyword arguments of select_n_clusters go to the model, and model selects another class (model=GK, for instance).

Notes

Indices disagree, and none is reliable alone. Fukuyama-Sugeno in particular tends to tie neighboring values of \(c\) on well separated data. Compare a few and look at the scores, not only at the winner.