This document compares the seriation methods available in the package
seriation using a sample of 30 flowers from the popular
Iris dataset and randomizes the order of the objects.
We first register more seriation methods. Some of these methods require installing additional packages.
The following methods will be used (a few slow methods are skipped).
methods <- sort(list_seriation_methods("dist"))
methods <- setdiff(methods, c("BBURCG", "BBWRCG", "Enumerate", "GSA", "SGD", "SGLS"))
methods ## [1] "ARSA" "DendSer" "DendSer_ARc" "DendSer_BAR"
## [5] "DendSer_LPL" "DendSer_PL" "GW" "GW_average"
## [9] "GW_complete" "GW_single" "GW_ward" "HC"
## [13] "HC_average" "HC_complete" "HC_single" "HC_ward"
## [17] "Identity" "MDS" "MDS_angle" "MDS_smacof"
## [21] "OLO" "OLO_average" "OLO_complete" "OLO_single"
## [25] "OLO_ward" "QAP_2SUM" "QAP_BAR" "QAP_Inertia"
## [29] "QAP_LS" "R2E" "Random" "Reverse"
## [33] "SPIN_NH" "SPIN_STS" "Sammon_mapping" "Spectral"
## [37] "Spectral_norm" "TSP" "VAT" "isoMDS"
## [41] "optics"
Details about the method can be found in the manual
page for seriate().
We use a loop to run the function seriate() with each
method and calculate criterion measures which indicate how good the
order is.
orders <- list()
criterion <- list()
for (m in methods) {
cat(m)
tm <- system.time(orders[[m]] <- seriate(d, method = m))
criterion[[m]] <- data.frame(time = tm[1]+tm[2], rbind(criterion(d, orders[[m]])))
cat(" took", tm[1]+tm[2], "sec.\n")
}## ARSA took 0.376 sec.
## DendSer took 0.727 sec.
## DendSer_ARc took 1.018 sec.
## DendSer_BAR took 0.515 sec.
## DendSer_LPL took 0.577 sec.
## DendSer_PL took 0.505 sec.
## GW took 0.251 sec.
## GW_average took 0.242 sec.
## GW_complete took 0.258 sec.
## GW_single took 0.249 sec.
## GW_ward took 0.249 sec.
## HC took 0.227 sec.
## HC_average took 0.208 sec.
## HC_complete took 0.329 sec.
## HC_single took 0.241 sec.
## HC_ward took 0.264 sec.
## Identity took 0.264 sec.
## MDS took 0.258 sec.
## MDS_angle took 0.235 sec.
## MDS_smacof took 0.228 sec.
## OLO took 0.206 sec.
## OLO_average took 0.216 sec.
## OLO_complete took 0.203 sec.
## OLO_single took 0.204 sec.
## OLO_ward took 0.2 sec.
## QAP_2SUM took 0.199 sec.
## QAP_BAR took 0.191 sec.
## QAP_Inertia took 0.197 sec.
## QAP_LS took 0.185 sec.
## R2E took 0.187 sec.
## Random took 0.183 sec.
## Reverse took 0.192 sec.
## SPIN_NH took 0.202 sec.
## SPIN_STS took 0.198 sec.
## Sammon_mapping took 0.186 sec.
## Spectral took 0.183 sec.
## Spectral_norm took 0.183 sec.
## TSP took 0.201 sec.
## VAT took 0.197 sec.
## isoMDS took 0.213 sec.
## optics took 0.194 sec.
We align the seriation orders. The reason is that an order 1, 2, 3
and 3, 2, 1 are equivalent and just an artifact of the algorithm.
Aligning will reverse some orders so they are better aligned. Then we
sort the orders from best to worst according to a popular seriation
criterion measure called Gradient_weighted.
orders <- ser_align(orders)
best_to_worse <- order(criterion[["Gradient_weighted"]], decreasing = TRUE)
orders <- orders[best_to_worse]
criterion <- criterion[best_to_worse, ]We can compare the seriation methods by how similar the orders are that they produce (measured using Spearman). The following code calculates distances between orders and then performs hierarchical clustering.
The reordered dendrogram clearly shows a group of methods based on hierarchical clustering focused on path length and another group that tries to optimize the other seriation measures.
Here is a table to compare the seriation methods on different
criterion measures. Use the interactive table to sort the methods given
different measures. Note that some are maximized and some should be
minimized. Details about the measures can be found in the manual
page for criterion().
Plot the reordered dissimilarity matrices. Dark blocks along the main diagonal mean that the order reveals a “cluster” of similar objects. The Iris dataset contains three species, but two of them are very similar, so we expect to see one smaller block and one larger block.
Matrix seriation reorders rows and columns of a data matrix. We perform the same steps as for distances in the previous section.
methods <- sort(list_seriation_methods("matrix"))
# AOE if for correlation matrices only
methods <- setdiff(methods, c("AOE"))
methods ## [1] "BEA" "BEA_TSP" "BK_unconstrained" "CA"
## [5] "Heatmap" "Identity" "LLE" "Mean"
## [9] "PCA" "PCA_angle" "Random" "Reverse"
Performing seriation.
orders <- list()
criterion <- list()
for (m in methods) {
cat(m)
tm <- system.time(orders[[m]] <- seriate(x, method = m))
criterion[[m]] <- data.frame(time = tm[1]+tm[2], rbind(criterion(x, orders[[m]])))
cat(" took", tm[1]+tm[2], "sec.\n")
}## BEA took 0.6 sec.
## BEA_TSP took 0.595 sec.
## BK_unconstrained took 0.211 sec.
## CA took 0.243 sec.
## Heatmap took 0.681 sec.
## Identity took 0.231 sec.
## LLE took 0.226 sec.
## Mean took 0.214 sec.
## PCA took 0.217 sec.
## PCA_angle took 0.231 sec.
## Random took 0.259 sec.
## Reverse took 0.271 sec.
best_to_worse <- order(criterion[["Moore_stress"]], decreasing = FALSE)
orders <- orders[best_to_worse]
criterion <- criterion[best_to_worse, ]