---
title: "Extraction"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{a01_extraction}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
```

## Overview

Phase 1 derives stratified prevalence tables from the records in an OMOP CDM database. The main entry point is `extract_all()`, which runs the full pipeline for one or more clinical domains.

Each domain produces 4 tables:

| Table | Content | Key columns |
|----|----|----|
| `*_prevalence` | Patient counts per concept x year x sex x age group | concept_id, year, sex, age_group, patient_count, denominator, prevalence |
| `*_info` | Concept metadata | concept_id, concept_name, vocabulary_id, n_patients_total |
| `*_chapters` | Hierarchy-based chapter assignments | concept_id, chapter_type, chapter_id, chapter_name |
| `*_attributes` | SNOMED relationship targets | concept_id, relationship, target_concept_id, target_concept_name |

Plus two shared tables: `demographics` (birth year x sex) and `death_counts` (deaths by stratum).

## Connecting to a database

Syrona connects through `CDMConnector`. A production OMOP CDM is normally a PostgreSQL database, so `syrona_connect_pg()` is the main entry point; `syrona_connect()` opens a local DuckDB file (for example an Eunomia or Synthea test dataset).

### PostgreSQL database

```{r, eval=FALSE}
library(syrona)

# Direct connection
db <- syrona_connect_pg(
  host = "db-server.example.com",
  dbname = "omop",
  user = "analyst",
  cdm_schema = "cdm",
  write_schema = "results_analyst"
)

# Via SSH tunnel (start tunnel first: ssh -L 5432:localhost:5432 user@server)
db <- syrona_connect_pg(
  host = "localhost",
  dbname = "omop",
  user = "analyst",
  cdm_schema = "ohdsi_cdm_202511",
  write_schema = "results_analyst"
)
```

**Credentials.** You can pass `password = "..."` directly, but it is cleaner to omit it and let RPostgres read the password from a `~/.pgpass` file or the `PGPASSWORD` environment variable — for example, set `PGPASSWORD=...` in your `~/.Renviron`. This keeps the password out of your R scripts and command history.

The `write_schema` parameter tells CDMConnector where it can create temporary tables. This is required for cohort operations and some extraction queries. On PostgreSQL, this is typically a user-specific results schema.

### DuckDB (local files)

```{r, eval=FALSE}
# Read-only (default) - safe for shared databases
db <- syrona_connect("path/to/omop.duckdb")

# Writable - needed if you want to create cohort tables in the same DB
db <- syrona_connect("path/to/omop.duckdb", read_only = FALSE)
```

## Running extraction

### All domains (default)

```{r, eval=FALSE}
tables <- extract_all("Dataset_A", db = db)
```

This extracts conditions, procedures, and drugs. Results are saved to `data/sources/Dataset_A/` and returned as a named list.

### Single domain

```{r, eval=FALSE}
# Conditions only (fastest)
tables <- extract_all("Dataset_A", db = db, domains = "conditions")

# Drugs only
tables <- extract_all("Dataset_A", db = db, domains = "drugs")

# Conditions + procedures (no drugs)
tables <- extract_all("Dataset_A", db = db, domains = c("conditions", "procedures"))
```

### Shorthand: pass a DuckDB path directly

```{r, eval=FALSE}
# This connects, extracts, and disconnects automatically
tables <- extract_all("Synthetic", db = "path/to/omop.duckdb")
```

## What each extractor does

### Denominators (ACHILLES-116)

`extract_denominators()` computes the number of persons observed per year x sex x age group. This is the denominator for all prevalence calculations. A person is counted in a year if their observation period overlaps that year.

Age groups are 10-year decades (0-9, 10-19, ..., 70-79, 80+). Ages above 80 are clamped into a single group for statistical stability.

### Condition prevalence (ACHILLES-404)

`extract_condition_prevalence()` counts persons with at least one condition occurrence per concept x year x sex x age group. Events must fall within the person's observation period. Only standard SNOMED concepts (concept_id != 0) are included.

### Condition chapters

`extract_condition_chapters()` assigns each condition concept to chapters via three classification systems:

-   **body_system** - SNOMED hierarchy from "Disorder of body system" (L1 + L2)
-   **disease_category** - SNOMED hierarchy from "Disease" (L1 + L2)
-   **icd10_chapter** - ICD-10 chapters via SNOMED-to-ICD-10 mapping (code prefix matching)

A concept can belong to multiple chapters. Concepts without an ICD-10 mapping get an "(Unmapped)" pseudo-chapter.

### Drug prevalence

`extract_drug_prevalence()` rolls up drug exposures to the **Ingredient** level via `concept_ancestor`. This means a prescription for "Aspirin 100mg tablet" counts toward the "Aspirin" ingredient. One row per ingredient x year x sex x age group.

### Drug chapters

`extract_drug_chapters()` assigns each ingredient to **ATC 1st level** chapters (e.g. "A. Alimentary tract and metabolism") via `concept_ancestor`.

### Procedure chapters

`extract_procedure_chapters()` assigns procedures to two SNOMED hierarchies:

-   **by_method** - from "Procedure by method" (L1 + L2)
-   **by_site** - from "Procedure by site" (L1 + L2)

## k-Anonymity

After extraction, `apply_k_anonymity()` suppresses small-cell counts (default k=5):

1.  Concepts with fewer than k total patients are dropped entirely
2.  Individual prevalence rows (strata) with fewer than k patients are suppressed
3.  Concepts that lose all prevalence rows to suppression are moved to a `*_rare` table (keeps aggregate counts but no stratified data)

This ensures no individual can be identified from the output tables.

## Loading and listing datasets

```{r, eval=FALSE}
# List all extracted datasets
list_datasets()
#> [1] "Dataset_A" "Dataset_B" "Synthetic"

# Load a previously extracted dataset
d <- load_dataset("Dataset_A")
names(d)
#> [1] "condition_prevalence" "condition_info" "condition_chapters" ...
```

## Disconnect

```{r, eval=FALSE}
syrona_disconnect(db)
```
