--- title: "Using dbscan from Python" author: "Michael Hahsler" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Using dbscan from Python} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} ### run with: #DBSCAN_RUN_PYTHON_VIGNETTE=true Rscript -e 'rmarkdown::render("vignettes/python.Rmd")' run_python <- identical( tolower(Sys.getenv("DBSCAN_RUN_PYTHON_VIGNETTE")), "true" ) knitr::opts_chunk$set(collapse = TRUE, comment = "#>") if (run_python) { Sys.setenv(RETICULATE_PYTHON = "managed") reticulate::py_require(c("numpy", "pandas", "rpy2")) } ``` R, the R package **dbscan**, and the Python package **rpy2** need to be installed. The following example reads the iris data and calls the R implementation of DBSCAN from Python. The Python chunks are not run during a normal package build. To execute them while rendering this vignette, set the environment variable `DBSCAN_RUN_PYTHON_VIGNETTE=true`. The first run uses **reticulate** to create a managed Python environment and install **numpy**, **pandas**, and **rpy2**. The environment is cached for later use, and all Python chunks run in one shared session. ```{python, eval=run_python} import pandas as pd import numpy as np import rpy2.robjects as ro # Prepare data. iris = pd.read_csv( "https://archive.ics.uci.edu/ml/machine-learning-databases/iris/iris.data", header=None, names=["SepalLength", "SepalWidth", "PetalLength", "PetalWidth", "Species"], ) iris_numeric = iris[["SepalLength", "SepalWidth", "PetalLength", "PetalWidth"]] # Import the R dbscan package. from rpy2.robjects import packages from rpy2.robjects import pandas2ri dbscan = packages.importr("dbscan") # Convert the pandas data frame using local conversion rules. conversion_rules = ro.default_converter + pandas2ri.converter with conversion_rules.context(): iris_r = ro.conversion.get_conversion().py2rpy(iris_numeric) db = dbscan.dbscan(iris_r, eps=0.5, minPts=5) print(db) ``` Extract the cluster assignment vector as a NumPy array: ```{python, eval=run_python} labels = np.asarray(db.rx2("cluster"), dtype=int) labels ``` Cluster label `0` identifies noise; positive integers identify clusters.