DARTER pitfalls
Quirks and known issues specific to project 708421
This page supplements the general DST pitfalls with issues specific to the DARTER project.
1. Check that parquet files are up to date
Most registers are as of 2026 updated to end of 2024 (confirmed by Anders Aasted Isaksen/Marie Kempf Frydendahl, DARTER team).
# Check when the parquet folder was last updated:
file.info("E:/workdata/708421/cleaned-data/parquet-registers/bef/")$mtimeDeaths are a special case, and it is not about the update date. dodsaars is a closed register: it ends in 2001 and no update will ever extend it. Deaths belong in dod (Døde i Danmark), which runs to 2025. On DARTER dod is not in cleaned-data, so see Register paths and datastores for how to read it, and pitfall 1 - four death registers for the general version.
The same may apply to other registers. Always confirm that coverage matches your study period before running the pipeline.
If the parquet file does not cover your study period: You need to extract data from the raw SAS file on DST. Contact your data manager - they can help with raw data access and conversion.
library(fastreg) # read_register(): opens a parquet register by name
library(dplyr) # the verbs below, and the pipe
# Any register: check that its coverage reaches the end of your follow-up
akm <- read_register("akm") %>% rename_with(tolower) # lazy connection
akm %>% summarise(min(year), max(year)) %>% collect()Consequence of missing coverage: comparators and BS patients whose events fall after the parquet file’s end date look like nothing happened to them - this affects censoring and matching in 01_build_cohorts.R, with no error message.
2. Surgery and procedures
Procedure codes are split across two registers by period:
lpr_sksopr(parquet-registers) - procedures and surgery 1996–2018, joined tolpr_admviarecnumprocedurer_kirurgi(parquet-external) - 2019 and onwards, joined tolpr_a_kontaktviadw_ek_forloeb
dw_ek_kontakt is NA for all rows in procedurer_kirurgi (confirmed 2026-06-02). Use dw_ek_forloeb - not dw_ek_kontakt - to fetch pnr from lpr_a_kontakt.
There is a third table: lpr_a_procregistrering. It is the general LPR3 procedure table, it exists on DARTER, and its dw_ek_kontakt is reported to be populated for the large majority of rows - so it joins to lpr_a_kontakt the same way lpr_a_diagnose does, and avoids the problem above entirely. The column names differ (proc_kode, proc_starttidspunkt), so the code is not interchangeable. See Overview of registers for a code example and the checks to run before you rely on it.
# WRONG - dw_ek_kontakt is NA:
proc %>% left_join(contacts, by = "dw_ek_kontakt") # joins nothing
# CORRECT - use dw_ek_forloeb:
proc <- read_register("procedurer_kirurgi") %>%
rename_with(tolower) %>%
left_join(
read_register("lpr_a_kontakt") %>%
rename_with(tolower) %>%
select(dw_ek_forloeb, pnr),
by = "dw_ek_forloeb"
)3. lpr_a_diagnose - “a” does not mean A-type diagnoses
The table is called lpr_a_diagnose - “a” refers to the analysis model designation (LPR_A series). It is not a filter on A-type diagnoses. The table contains A, B and G. You still need to filter on diag_kode_type.
4. nmi_count ≠ nmi_score
| Variable | What it is |
|---|---|
nmi_score |
Weighted score - Nordic Multimorbidity Index (50 predictors with individual weights) |
nmi_count |
Simple count of the number of chronic conditions (33 possible) |
If you use nmi_count in your Cox model instead of nmi_score, you are adjusting for something different than you think.
5. LPR3 - filter on lprindberetningssystem == "LPR3"
Besides the actual LPR3 reports, lpr_a_kontakt also contains older data that is already present in LPR2 (lpr_adm), loaded into the new LPR_A table. If you combine LPR2 and LPR3 without filtering, the same contacts are counted twice - you get duplicated rows. lprindberetningssystem == "LPR3" keeps only the rows from the LPR3 system - the fix Anders Aasted Isaksen (DARTER team, 2026) originally suggested for this duplication problem.
The same column is on lpr_a_diagnose. Filter that table too - for the same reason as the contact table: an unfiltered diagnosis table still carries the same LPR2-reload overlap, and double-counts the same way.
# CORRECT - filter both tables to LPR3-system rows:
lpr3_k <- read_register("lpr_a_kontakt") %>%
rename_with(tolower) %>%
filter(lprindberetningssystem == "LPR3") # keep only rows from the LPR3 system - removes overlapping rows
lpr3_d <- read_register("lpr_a_diagnose") %>%
rename_with(tolower) %>%
filter(lprindberetningssystem == "LPR3") # same column on the diagnosis tableget_lpr_diagnoses() in darter-index.qmd is updated with this filter. If you have copies of LPR3 code in your own scripts, you must add it manually.
On DARTER, this filter currently drops a real quarter of contacts - not just duplicates. DST’s LPR2 register closed 31 March 2019, but DARTER’s lpr_adm/lpr_diag delivery available under cleaned_data currently stops 31 December 2018, a full quarter earlier (confirmed directly, not DST’s documented date - see DARTER register paths). The LPR2-tagged rows still sitting in lpr_a_kontakt for January-March 2019 are the only record of those contacts anywhere in this delivery: filtering to lprindberetningssystem == "LPR3" removes them entirely, not just de-duplicates them. Direct measurement confirms the loss is bounded to that one quarter: almost nothing (under 100 rows, out of tens of millions) carries an LPR2-tagged start date after 31 March 2019, so this is not an open-ended problem spanning years.
Check your own lpr_adm’s last date before trusting the simple filter above - if it has moved to 31 March 2019 or later, you are done, the simple filter is safe as written. If it has not, do not just drop the lprindberetningssystem filter to get the missing quarter back - that also brings back the real 2017-2018 duplicates the filter exists to remove. The fix is a date cutover, not removing the filter: keep genuine LPR3 rows, and keep LPR2-tagged rows only for the window your own lpr_adm does not cover yet.
# 1. Where does DARTER's own LPR2 delivery actually stop? This decides which
# LPR2-tagged rows in lpr_a_kontakt are genuine duplicates (already in
# lpr_adm) versus genuinely missing from lpr_adm (the gap to recover).
lpr_adm_last_date <- read_register("lpr_adm") %>%
rename_with(tolower) %>%
summarise(max_date = max(d_inddto, na.rm = TRUE)) %>%
collect() %>%
pull(max_date)
# 2. Contacts: read_register() does not resolve lpr_a_kontakt on DARTER yet -
# use read_parquet_dataset() with the full path as a temporary workaround
# (see Parquet and fastreg). This is expected to be fixed soon; once
# read_register("lpr_a_kontakt") resolves, switch back to it. Drop all of
# LPR1 and MiniPAS (private hospitals), then drop only the LPR2-tagged rows
# lpr_adm already covers. What survives is every genuine LPR3 contact, plus
# the LPR2-tagged contacts for the gap lpr_adm is missing.
lpr3_k <- read_parquet_dataset("E:/workdata/708421/cleaned-data/parquet-registers/lpr_a_kontakt") %>%
rename_with(tolower) %>%
filter(lprindberetningssystem %in% c("LPR2", "LPR3")) %>% # only the two systems that can duplicate lpr_adm
mutate(start_date = as.Date(kont_starttidspunkt)) %>%
filter(!(lprindberetningssystem == "LPR2" & start_date <= lpr_adm_last_date)) # drop only the LPR2 rows lpr_adm already has
# 3. Diagnoses: same temporary read_parquet_dataset() workaround as step 2 -
# switch back to read_register("lpr_a_diagnose") once it resolves on DARTER.
# This table carries no date of its own (unlike kontakt), so it cannot
# repeat step 2's date logic directly. Instead, keep only the diagnosis
# rows that belong to a contact step 2 already decided to keep, matched on
# dw_ek_kontakt. Filtering this table to "LPR3" only would silently drop
# the diagnosis codes for the recovered gap-window LPR2 contacts, making
# them look like contacts with no diagnosis at all - that is the actual
# reason for the semi_join below.
lpr3_d <- read_parquet_dataset("E:/workdata/708421/cleaned-data/parquet-registers/lpr_a_diagnose") %>%
rename_with(tolower) %>%
filter(lprindberetningssystem %in% c("LPR2", "LPR3")) %>%
semi_join(
lpr3_k %>% select(dw_ek_kontakt),
by = "dw_ek_kontakt"
)The 2017-2018 LPR2-tagged rows stay excluded in both tables (their start_date is before lpr_adm_last_date, so step 2’s filter() removes them, and step 3 only keeps what step 2 kept) - only the January-March 2019 gap gets pulled back in, diagnosis codes included. This is also self-correcting: once Marie’s delivery update lands and lpr_adm_last_date reaches 31 March 2019, step 2’s LPR2 branch stops matching anything and both tables quietly become equivalent to lprindberetningssystem == "LPR3" again - no need to remember to remove this code later.
Separately: private hospitals are not in lpr_adm/lpr_diag at all - before 2019 they reported to a different register, MiniPas - LPR, which must be requested separately and is not recovered by filtering LPR2 or LPR3 harder. See the fuller caveat in Extract from LPR before relying on counts that span 2019.
6. In the DARTER delivery, FAIK carries pnr from 2022
FAIK is documented by DST as one row per family per year, keyed on familie_id, which is why the standard recipe joins income to people via familie_id from BEF. In the DARTER delivery, the 2022 data onward also carries pnr, and the family’s row is repeated once per family member (confirmed by the DARTER team, August 2026).
This is a fact about this delivery, not about the register: neither DST’s variable list for FAIK nor DST’s order list (checked 2026-09-30) contains a pnr, for any year. Whether it comes from how the data was ordered or from how it was converted is not known, so another project may or may not see it.
If you join on familie_id without removing the copies, each person gets one row per family member for those years. Nothing errors. A single-year join just grows, and in the 3-year income mean each person is still averaged on their own, but a year from 2022 onward counts once per family member: n_years can exceed 3, and a window such as 2020-2022 gives 2022 more weight than the two years before it.
# Check before you join - more than one row per family and year means it applies
faik %>%
count(familie_id, year) %>%
filter(n > 1) %>%
count() %>%
collect()
# distinct() is only a safe fix if every copy carries the same income - expect 0
faik %>%
group_by(familie_id, year) %>%
summarise(n_values = n_distinct(famaekvivadisp_13), .groups = "drop") %>%
filter(n_values > 1) %>%
count() %>%
collect()
# Fix: collapse back to one row per household-year before joining
faik <- faik %>%
distinct(familie_id, year, famaekvivadisp_13)Full explanation and the alternative (joining on pnr directly) is in Socioeconomic variables.
7. Laboratory results - use laboratorieproevesvar_
The new laboratory data register is called laboratorieproevesvar_ and contains >2.2 billion rows. The old lab_forsker / lab_dm_forsker still exists but covers the same data - use only one source to avoid duplicates.
lab <- read_register("laboratorieproevesvar") %>%
rename_with(tolower) %>%
rename(pnr = cprnummer) %>% # the person column is cprnummer here, patient_cpr in the other two tables
semi_join(tibble(pnr = cohort$pnr), by = "pnr", copy = TRUE) %>% # filter BEFORE collect - the register is very large
select(pnr, analysiscode, samplingdate, samplevalue) %>% # NPU is the coding system, analysiscode is the column
collect()
# samplevalue is character - can contain "not detected", "negative" etc.See also
- General DST pitfalls: 13 pitfalls that apply to all projects
- Register paths and datastores: confirmed paths and access methods