--- title: "Starting on S3" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Starting on S3} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", eval = FALSE ) ``` ```{=html} ```
**Goal:** Stand up a versioned datom project whose data lives in **Amazon S3**, and onboard a study's files with `datom_sync()`. The steps are the ones from [Getting Started](getting-started.html) -- one file, an update, a no-op, a batch -- and the only difference is the store you build.
> **Want to try datom locally first?** Start with > [Getting Started](getting-started.html). It uses a folder on your machine > instead of S3 and needs no AWS account. The functions are the same. You look after the data for **study001**, a clinical trial, and your team already works in S3. You want the first extract to land in the shared bucket, versioned from day one. datom keeps the data in S3 and the record of every version in a git repository, so the history can be read and reproduced from any machine.
**Two locations, two roles:** - **Input folder** -- where new files land before datom takes them in. Files here can be overwritten or deleted. Syncing a file datom already holds does nothing. - **Storage** -- once `datom_sync()` takes a file in, S3 holds a versioned copy. The input file is no longer needed.
## Requirements - **A GitHub account** with a personal access token (PAT) scoped to `repo`. datom creates the metadata repository through the GitHub API with it; the `gh` CLI is not needed. - **An S3 bucket** you can read and write. datom does not create buckets: encryption, versioning and retention are your organization's policy. This article uses one bucket per study, with a folder per project inside it. - **AWS credentials** (an access key and a secret key) for that bucket. - **The `git2r`, `rio` and `keyring` packages.** datom uses `git2r` for the metadata repository and `rio` to read the files you sync; `keyring` holds your secrets in the OS keychain. Store the three secrets in your keychain once: ```{r keyring-setup} keyring::key_set(service = "GITHUB_PAT") keyring::key_set(service = "AWS_ACCESS_KEY_ID") keyring::key_set(service = "AWS_SECRET_ACCESS_KEY") ``` ## Load your secrets This is the only place the keychain is read. Everything after it reads the environment with `Sys.getenv()`, so the secrets never appear in your code. ```{r secrets} Sys.setenv( GITHUB_PAT = keyring::key_get(service = "GITHUB_PAT"), AWS_ACCESS_KEY_ID = keyring::key_get(service = "AWS_ACCESS_KEY_ID"), AWS_SECRET_ACCESS_KEY = keyring::key_get(service = "AWS_SECRET_ACCESS_KEY") ) ``` On CI or in a container, set these three environment variables through your platform's secret store and skip this chunk. ## Settings Every value used more than once is set here, so changing one means changing it in one place. ```{r settings} library(datom) # --- Settings you control ---------------------------------------------------- bucket <- "study001" # one bucket per study region <- "us-east-1" project_imported <- "study001-imported" # recorded in the project's metadata prefix_imported <- "imported/" # this project's folder in the bucket repo_imported <- "study001-imported" # GitHub repo name # Local working folder for the metadata repository. The data never lands here; # it goes straight to S3. workdir_imported <- fs::path(tempdir(), "study001-imported") ``` The project, repo and folder names match by convention only. They are separate settings because they do not have to match. ## Where the data goes ``` s3://study001/ imported/datom/ onboarded tables (this article) ``` `imported` is datom's word for a table that came in from a file. datom manages the `datom/` folder under each prefix: it reads and writes the data, metadata and version records there, and leaves anything else in the bucket alone. ## Build the store A **store** says where the data lives and how to reach it. Giving it a GitHub token makes it a **writer** store: it can create the metadata repository and record new versions. ```{r store-write} store_write_imported <- datom_store( data = datom_store_s3( bucket = bucket, prefix = prefix_imported, region = region, access_key = Sys.getenv("AWS_ACCESS_KEY_ID"), secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY") ), github_pat = Sys.getenv("GITHUB_PAT") ) ``` datom checks that it can reach the bucket as the store is built, so a wrong credential shows up here rather than at the first sync. ## Create the repository ```{r init} datom_init_repo( path = workdir_imported, project_name = project_imported, store = store_write_imported, create_repo = TRUE, repo_name = repo_imported ) #> v Created GitHub repo ".../study001-imported". #> v Initialized datom repository "study001-imported" at '.../study001-imported' ``` This creates the GitHub repository, clones it into `workdir_imported`, and commits a `project.yaml` recording where the project's data lives. No data goes to GitHub, only metadata. There is no `mode` argument here. Leaving it out gives an ordinary repository, one that takes files in with `datom_sync()`. The clone also gets an `input_files/` folder. git ignores it, so nothing placed there is ever committed. It is where `datom_sync()` looks for new files. ## Connect ```{r conn-write} conn_write_imported <- datom_get_conn( path = workdir_imported, store = store_write_imported ) print(conn_write_imported) #> #> -- datom connection #> * Project: "study001-imported" #> * Backend: "s3" #> * Role: "developer" #> * Data root: "study001" #> * Data prefix: "imported/" #> * Data region: "us-east-1" #> * Governance: not attached #> * Path: '.../study001-imported' #> * Data repo: ``` ## Step 1: Sync one file The month-1 extract of the demographics table has arrived. Put it in the input folder: ```{r step1-write} inputs_imported <- fs::path(workdir_imported, "input_files") write.csv( x = datom_example_data(domain = "dm", cutoff_date = "2026-01-28"), file = fs::path(inputs_imported, "dm.csv"), row.names = FALSE ) ``` Scan the folder, then sync: ```{r step1-sync} manifest <- datom_sync_manifest(conn = conn_write_imported) #> i Scanned 1 file: 1 new, 0 changed, 0 unchanged. synced <- datom_sync(conn = conn_write_imported, manifest = manifest) #> i Syncing 1 table... #> v Wrote "dm" (full): "153bff41" #> v "dm" synced (new). #> i Sync complete: 1 succeeded, 0 failed, 0 skipped. ``` The file was converted to parquet and uploaded to S3. The version record was committed to the metadata repository and pushed to GitHub. `datom_sync()` returns the manifest with a `result` for each file, kept here as `synced`. ```{r step1-list} datom_list(conn = conn_write_imported) #> name kind current_version current_data_sha last_updated #> 1 dm table 153bff41 decbafd2 2026-09-27T05:58:03Z dm_history <- datom_history(conn = conn_write_imported, name = "dm", short_hash = TRUE) dm_history[, c("version", "timestamp", "commit_message")] #> version timestamp commit_message #> 1 153bff41 2026-09-27T05:58:03Z Sync dm (new) ``` The input file is no longer needed. The table reads from S3: ```{r step1-delete} fs::file_delete(fs::path(inputs_imported, "dm.csv")) nrow(datom_read(conn = conn_write_imported, name = "dm")) #> [1] 4 ``` ## Step 2: Update one file The month-2 extract arrives with new subjects: ```{r step2} write.csv( x = datom_example_data(domain = "dm", cutoff_date = "2026-02-28"), file = fs::path(inputs_imported, "dm.csv"), row.names = FALSE ) manifest <- datom_sync_manifest(conn = conn_write_imported) #> i Scanned 1 file: 0 new, 1 changed, 0 unchanged. synced <- datom_sync(conn = conn_write_imported, manifest = manifest) #> i Syncing 1 table... #> v Wrote "dm" (full): "0fac26cd" #> v "dm" synced (changed). #> i Sync complete: 1 succeeded, 0 failed, 0 skipped. ``` Both versions stay readable. The newest is the default; an older one is read by its version. The first 8 characters of a version are enough, as with a git commit: ```{r step2-read} dm_history <- datom_history(conn = conn_write_imported, name = "dm", short_hash = TRUE) dm_history[, c("version", "timestamp", "commit_message")] #> version timestamp commit_message #> 1 0fac26cd 2026-09-27T05:58:10Z Sync dm (changed) #> 2 153bff41 2026-09-27T05:58:03Z Sync dm (new) nrow(datom_read(conn = conn_write_imported, name = "dm")) #> [1] 16 dm_version <- dm_history$version[nrow(dm_history)] # oldest row: month 1 nrow(datom_read(conn = conn_write_imported, name = "dm", version = dm_version)) #> [1] 4 ``` ## Step 3: Sync again with nothing new ```{r step3} manifest <- datom_sync_manifest(conn = conn_write_imported) #> i Scanned 1 file: 0 new, 0 changed, 1 unchanged. synced <- datom_sync(conn = conn_write_imported, manifest = manifest) #> i No new or changed files. Nothing to sync. ``` A file datom already holds is not uploaded again, so the sync is safe to run on a schedule. ## Step 4: A batch of files The month-3 extract brings four tables at once: demographics (`dm`), dosing (`ex`), labs (`lb`) and adverse events (`ae`). ```{r step4} for (domain in c("dm", "ex", "lb", "ae")) { write.csv( x = datom_example_data(domain = domain, cutoff_date = "2026-03-28"), file = fs::path(inputs_imported, paste0(domain, ".csv")), row.names = FALSE ) } manifest <- datom_sync_manifest(conn = conn_write_imported) #> i Scanned 4 files: 3 new, 1 changed, 0 unchanged. synced <- datom_sync(conn = conn_write_imported, manifest = manifest) #> i Syncing 4 tables... #> v Wrote "ae" (full): "075773e9" #> v "ae" synced (new). #> v Wrote "dm" (full): "773e6862" #> v "dm" synced (changed). #> v Wrote "ex" (full): "8dbcc9a7" #> v "ex" synced (new). #> v Wrote "lb" (full): "435bccb0" #> v "lb" synced (new). #> i Sync complete: 4 succeeded, 0 failed, 0 skipped. ``` All four tables are now versioned in S3: ```{r step4-list} datom_list(conn = conn_write_imported) #> name kind current_version current_data_sha last_updated #> 1 dm table 773e6862 e547f03d 2026-09-27T05:58:22Z #> 2 ae table 075773e9 d5f8dd5a 2026-09-27T05:58:17Z #> 3 ex table 8dbcc9a7 ab96afc3 2026-09-27T05:58:26Z #> 4 lb table 435bccb0 5d419c60 2026-09-27T05:58:31Z ``` ## Reading as a reader A colleague who only reads needs bucket credentials and nothing else: no GitHub token and no clone. A store without a token is a **reader** store. ```{r reader} store_read_imported <- datom_store( data = datom_store_s3( bucket = bucket, prefix = prefix_imported, region = region, access_key = Sys.getenv("AWS_ACCESS_KEY_ID"), secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY") ) ) conn_read_imported <- datom_get_conn( store = store_read_imported, project_name = project_imported ) nrow(datom_read(conn = conn_read_imported, name = "lb")) #> [1] 205 ``` The read goes straight to S3, not through GitHub. ## Where you are - Four tables are versioned in `s3://study001/imported/datom/`. - Their history is in a GitHub repository; the data itself never went there. - The same sync steps work on S3 as on a local folder. ## What next - **Keep going.** [Citing a Set of Tables](citable-sets.html) builds on this data: it collects these tables, and the ones you derive from them, into one versioned, citable set. Continue in the same R session and skip the teardown below for now. - **Stop here.** Run the [teardown](#teardown). ## Governance and migration come later Two things are deliberately left out of this article: - **Governance**: a shared register of projects, managed reader access and managed teardown. It is optional, provided by the companion package `datomanager`, and starting on S3 does not commit you to it. - **Migration**: moving an existing project's data from one store to another while keeping its history. That is also a `datomanager` workflow. You started on S3, so you do not need it now. ## Teardown Delete the project's storage first, then its repository: ```{r teardown-imported} datom_storage_delete_prefix(conn = conn_write_imported) datom_repo_delete(conn = conn_write_imported, confirm = project_imported) ``` `datom_storage_delete_prefix()` deletes everything under `imported/datom/`, and it does not ask first. The rest of the bucket is left alone, and so is the bucket itself. `datom_repo_delete()` deletes the GitHub repository and the local clone. It needs the project name as `confirm`, and it does not touch storage.