Supportive Periodontal Therapy Clinical Examination Data 1.0.0

File: <base>/README.md (9,596 bytes)
# Supportive Periodontal Therapy Clinical Examination Data

Version 1.0.0

Longitudinal clinical data from 883 patients enrolled in supportive periodontal
therapy at the Medi School of Dental Hygiene (MSDH), Bern, Switzerland, between
1985 and 2011, comprising 11,842 supportive periodontal therapy visits.

## 1. Overview

The dataset records patient-level indicators of periodontal health at three
stages: at the initial examination before active periodontal therapy, at
re-evaluation after active periodontal therapy, and at every subsequent
supportive periodontal therapy visit. Clinical variables comprise the number of
teeth, the number of sites at each residual probing depth category, bleeding on
probing, and the intervals between visits. Demographic variables comprise sex,
age, smoking status and relevant medical history.

Probing depths were recorded at six sites per tooth by dental hygiene students
and re-measured by experienced clinical instructors; student recordings were
corrected where they differed from the instructor's. Sites of 0 to 3 mm were not
recorded, and depths of 8 mm or more were not recorded separately from 7 mm.
Bleeding on probing was recorded at four sites per tooth.

The data underlie three published analyses: patient compliance with scheduled
visits [1], the relationship between bleeding on probing and smoking status [2],
and the derivation and validation of an algorithm that computes supportive
periodontal therapy intervals from the residual probing depth profile [3].

## 2. Files

| File | Rows | Content |
|---|---|---|
| `01_initial_periodontal_therapy_data.csv` | 883 | One row per patient: demographics, medical history, smoking status, disease severity, clinical findings before and after active periodontal therapy, and summary measures of the supportive phase |
| `02_supportive_periodontal_therapy_data.csv` | 11,842 | One row per supportive periodontal therapy visit: clinical findings, assigned and algorithm-based intervals, and the time elapsed since the preceding visit |
| `03_supplementary_data.csv` | 11,842 | One row per visit: variables derived from the primary data, including cumulative site counts, percentages, the intermediate steps of the interval algorithm, and adherence to the computed interval |
| `04_data_corrections.csv` | 63 | Every visit whose probing depth values differ from the source records, with the original value, the value used, and the nature of the edit |
| `dictionary.csv` | — | Data dictionary describing all 93 variables of the four data files |

Files 02 and 03 are joined on `pat_id` and `spt_id`; both join to file 01 on
`pat_id`.

Probing depth was recorded at six sites per tooth and bleeding on probing at
four. Probing depth percentages therefore use `n_teeth_spt * 6` as the
denominator and bleeding on probing percentages use `n_teeth_spt * 4`. Both
denominators are applied in the accompanying scripts.

## 3. Format

All data files are comma-separated values encoded in UTF-8 without a byte order
mark, formatted according to RFC 4180. Missing values are empty fields. The
files can be opened with any spreadsheet application or read directly in R,
Python or comparable environments.

```r
apt <- read.csv("01_initial_periodontal_therapy_data.csv")
spt <- read.csv("02_supportive_periodontal_therapy_data.csv")
sup <- read.csv("03_supplementary_data.csv")
```

```python
import pandas as pd
apt = pd.read_csv("01_initial_periodontal_therapy_data.csv")
spt = pd.read_csv("02_supportive_periodontal_therapy_data.csv")
sup = pd.read_csv("03_supplementary_data.csv")
```

## 4. Two things to know before using the data

**Manual edits during data preparation.** Sixty-three of the 11,842 visits
received manual attention when these data were first prepared, and this was
documented at the time. They fall into two groups, flagged by the `imputed` and
`corrected` columns in files 02 and 03, and listed individually in
`04_data_corrections.csv`.

In 47 visits **no probing depth measurement existed** and values were inserted,
in 41 cases by carrying forward the preceding visit. Those inserted values have
been removed: the probing depth counts and everything derived from them are
blank. The tooth count and the bleeding on probing count were measured and are
retained.

In 16 visits **a measurement existed but was judged implausible and replaced** —
for instance a visit recording 25 sites of 6 mm between visits recording 8 and 7.
The replacement is retained, because the published analyses rest on it. Both
values are given in `04_data_corrections.csv`.

No visits were removed. Users reproducing the published analyses can restore the
original state exactly from `04_data_corrections.csv`.

**Study populations.** Reference [1] used all 883 patients, and so do the
descriptive analyses and the derivation of the interval algorithm in reference
[3]. Reference [2], the validation of the algorithm, and the mixed-effects model
of Table 2 in reference [3] used a subsample of 445 patients: those who were
systemically healthy at the initial examination and who attended supportive
periodontal therapy for at least five years. The subsample is flagged by
`subsample_5years` in file 01 and can equally be derived as
`med_history == "healthy"` and `spt_duration_years >= 5`.

That Table 2 rests on the subsample rather than on all 883 patients is not
stated in the publication. It follows from the covariate for disease severity,
which was recorded only for those 445 patients, and it is confirmed by the scale
of the published estimates; the header of `02_linear_mixed_effects_model.R` gives
the evidence.

## 5. Accompanying code

Scripts are provided in R and, for the two deterministic analyses, in Python as
well. They run on the data files in this dataset and require no other input.

| Script | Purpose | Requires |
|---|---|---|
| `01_compute_spt_algorithm.R` / `.py` | Reproduces every derived variable of file 03 from the primary data, including the interval algorithm | base R / pandas, numpy |
| `02_linear_mixed_effects_model.R` / `.py` | Reconstruction of the linear mixed-effects model reported in Table 2 of reference [3] | `lme4`, `lmerTest` / pandas, numpy |
| `03_reproduce_figure_2.R` / `.py` | Reproduces the empirically determined probing depth stability thresholds shown in Figure 2 of reference [3], and reports a sensitivity analysis | base R / pandas, numpy |

Each pair was written against the same specification and the two implementations
agree. Scripts 01 and 03 verify themselves against the published files when run
and print the result: both reproduce all 11,842 rows of the derived data and all
twenty published stability thresholds exactly.

Script 02 is a reconstruction, not a rerun: the code written for the 2019
analysis has not survived. Its header sets out what the reconstruction rests on
and how closely it comes to the published table. Two points there matter to
anyone reusing these data. Table 2 was fitted on the five-year subsample of 445
patients, not on all 883, because the disease severity covariate was recorded
only for that subsample. And the F column of Table 2 reports sequential (type I)
tests while the confidence intervals beside it are marginal; the script prints
both, and the marginal tests are the ones to build on.

The Python version of script 02 fits the random-intercept model directly from
the profiled REML criterion rather than through a modelling library, so that it
returns the same estimates, standard errors and F values as the R version.
Satterthwaite degrees of freedom are not implemented, so it prints no p values;
use the R version when p values are needed.

## 6. De-identification

The dataset contains no directly or indirectly identifying information. Patient
identifiers are sequential study numbers that cannot be linked back to clinical
records. All calendar dates have been removed; only the year of each visit is
retained. Ages above 89 years are aggregated to 90 in accordance with the HIPAA
Safe Harbor standard.

## 7. Ethics

The Cantonal Ethics Committee of Bern, Switzerland, determined in its decision of
26 March 2025 (BASEC Req-2025-00339) that the publication of this dataset does
not fall within the scope of the Swiss Human Research Act and that approval by an
ethics committee is not required. Permission to conduct the original studies was
granted by the Medi School of Dental Hygiene, Bern, in 2011.

## 8. Licence and citation

Released under the Creative Commons Attribution 4.0 International Public License.

When using this resource, please cite the PhysioNet project page and the original
publication [3].

## 9. Contact

Christoph A. Ramseier, Department of Periodontology, School of Dental Medicine,
University of Bern, Switzerland — christoph.ramseier@unibe.ch

## 10. References

1. Ramseier CA, Kobrehel S, Staub P, Sculean A, Lang NP, Salvi GE. Compliance of
   cigarette smokers with scheduled visits for supportive periodontal therapy.
   J Clin Periodontol. 2014;41(5):473-480. doi:10.1111/jcpe.12242

2. Ramseier CA, Mirra D, Schütz C, Sculean A, Lang NP, Walter C, Salvi GE.
   Bleeding on probing as it relates to smoking status in patients enrolled in
   supportive periodontal therapy for at least 5 years. J Clin Periodontol.
   2015;42:150-159. doi:10.1111/jcpe.12344

3. Ramseier CA, Nydegger M, Walter C, Fischer G, Sculean A, Lang NP, Salvi GE.
   Time between recall visits and residual probing depths predict long-term
   stability in patients enrolled in supportive periodontal therapy.
   J Clin Periodontol. 2019;46(2):218-230. doi:10.1111/jcpe.13041