Phase 1 derives stratified prevalence tables from the records in an
OMOP CDM database. The main entry point is extract_all(),
which runs the full pipeline for one or more clinical domains.
Each domain produces 4 tables:
| Table | Content | Key columns |
|---|---|---|
*_prevalence |
Patient counts per concept x year x sex x age group | concept_id, year, sex, age_group, patient_count, denominator, prevalence |
*_info |
Concept metadata | concept_id, concept_name, vocabulary_id, n_patients_total |
*_chapters |
Hierarchy-based chapter assignments | concept_id, chapter_type, chapter_id, chapter_name |
*_attributes |
SNOMED relationship targets | concept_id, relationship, target_concept_id, target_concept_name |
Plus two shared tables: demographics (birth year x sex)
and death_counts (deaths by stratum).
Syrona connects through CDMConnector. A production OMOP
CDM is normally a PostgreSQL database, so
syrona_connect_pg() is the main entry point;
syrona_connect() opens a local DuckDB file (for example an
Eunomia or Synthea test dataset).
library(syrona)
# Direct connection
db <- syrona_connect_pg(
host = "db-server.example.com",
dbname = "omop",
user = "analyst",
cdm_schema = "cdm",
write_schema = "results_analyst"
)
# Via SSH tunnel (start tunnel first: ssh -L 5432:localhost:5432 user@server)
db <- syrona_connect_pg(
host = "localhost",
dbname = "omop",
user = "analyst",
cdm_schema = "ohdsi_cdm_202511",
write_schema = "results_analyst"
)Credentials. You can pass
password = "..." directly, but it is cleaner to omit it and
let RPostgres read the password from a ~/.pgpass file or
the PGPASSWORD environment variable — for example, set
PGPASSWORD=... in your ~/.Renviron. This keeps
the password out of your R scripts and command history.
The write_schema parameter tells CDMConnector where it
can create temporary tables. This is required for cohort operations and
some extraction queries. On PostgreSQL, this is typically a
user-specific results schema.
This extracts conditions, procedures, and drugs. Results are saved to
data/sources/Dataset_A/ and returned as a named list.
extract_denominators() computes the number of persons
observed per year x sex x age group. This is the denominator for all
prevalence calculations. A person is counted in a year if their
observation period overlaps that year.
Age groups are 10-year decades (0-9, 10-19, …, 70-79, 80+). Ages above 80 are clamped into a single group for statistical stability.
extract_condition_prevalence() counts persons with at
least one condition occurrence per concept x year x sex x age group.
Events must fall within the person’s observation period. Only standard
SNOMED concepts (concept_id != 0) are included.
extract_condition_chapters() assigns each condition
concept to chapters via three classification systems:
A concept can belong to multiple chapters. Concepts without an ICD-10 mapping get an “(Unmapped)” pseudo-chapter.
extract_drug_prevalence() rolls up drug exposures to the
Ingredient level via concept_ancestor.
This means a prescription for “Aspirin 100mg tablet” counts toward the
“Aspirin” ingredient. One row per ingredient x year x sex x age
group.
extract_drug_chapters() assigns each ingredient to
ATC 1st level chapters (e.g. “A. Alimentary tract and
metabolism”) via concept_ancestor.
extract_procedure_chapters() assigns procedures to two
SNOMED hierarchies:
After extraction, apply_k_anonymity() suppresses
small-cell counts (default k=5):
*_rare table (keeps aggregate counts but no stratified
data)This ensures no individual can be identified from the output tables.