--- title: "Using dbscan with tidyverse" author: "Michael Hahsler" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Using dbscan with tidyverse} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` The **dbscan** package provides `tidy()`, `augment()`, and `glance()` methods for its clustering algorithms, making them easy to use with tidyverse, ggplot2, and [tidymodels](https://www.tidymodels.org/learn/statistics/k-means/). Load the packages and prepare the numeric variables from the iris data: ```{r data, message=FALSE, warning=FALSE} library(dbscan) library(tidyverse) x <- iris[, 1:4] db <- x %>% dbscan(eps = .42, minPts = 5) ``` Get cluster statistics as a tibble: ```{r tidy} tidy(db) ``` Visualize the clustering with ggplot2, using an x for noise points: ```{r plot, fig.alt="DBSCAN clusters in the iris data"} augment(db, x) %>% ggplot(aes(x = Petal.Length, y = Petal.Width)) + geom_point(aes(color = .cluster, shape = noise)) + scale_shape_manual(values = c(19, 4)) ```