Get started

library(unsum)

Here is a brief walkthrough of using CLOSURE in unsum.

Call closure_generate() to run the CLOSURE algorithm. Enter mean, SD, and sample size that you read in a paper. For scale_min and scale_max, use the empirical minimum and maximum if available. Otherwise, use the more fundamental scale bounds, e.g., 1 and 7 for a 1-7 scale.

The mean and sd arguments must be strings to preserve trailing zeros. Note that CLOSURE can only be used if the values must be integers: e.g., a value can be 2 or 3, but not 2.5.

data <- closure_generate(
  mean = "3.5",
  sd = "1.8",
  n = 80,
  scale_min = 1,
  scale_max = 5
)
#> 
#> ✔ All CLOSURE results found

First create a plot of the mean sample found by CLOSURE. This gives us a sense of the overall results, which are quite polarized:

closure_plot_bar(data)

Barplot of `data`, the CLOSURE output. It specifically visualizes the `f_average` column of the `frequency` tibble, but also gives percentage figures, similar to the `f_relative` column. The overall shape is a strongly polarized distribution.

You can customize the plot, e.g., to show the sum of all samples found instead of the average sample, or only percentages, or different colors. See documentation at closure_plot_bar(). However, the default should be informative enough for a start.

CLOSURE results

Now let’s look at the results themselves:

data
#> 
#> ── CLOSURE results: 2,215 samples ──────────────────────────────────────────────
#> $inputs · how `closure_generate()` produced these data
#> # A tibble: 1 × 8
#>   technique mean  sd        n scale_min scale_max rounding   threshold
#>   <chr>     <chr> <chr> <dbl>     <dbl>     <dbl> <chr>          <dbl>
#> 1 CLOSURE   3.5   1.8      80         1         5 up_or_down         5
#> $metrics_main · key statistics about the generated samples
#> # A tibble: 1 × 2
#>   samples_all values_all
#>         <dbl>      <dbl>
#> 1        2215     177200
#> $metrics_horns · on the distribution of horns index values
#> # A tibble: 1 × 9
#>    mean uniform     sd     cv    mad   min median   max  range
#>   <dbl>   <dbl>  <dbl>  <dbl>  <dbl> <dbl>  <dbl> <dbl>  <dbl>
#> 1 0.792     0.5 0.0251 0.0317 0.0189 0.756  0.787 0.844 0.0877
#> $directory · where the results are saved on disk
#> # A tibble: 1 × 1
#>   path 
#>   <chr>
#> 1 <NA>
#> With hidden elements:
#> ℹ Access $modality_counts for min/max counts per scale value (5 rows)
#> ℹ Access $modality_pairs for frequency ordering between adjacent values (4 rows)
#> ℹ Access $modality_conclusion for flagging modality and J-shape status (1 row)
#> ℹ Access $frequency for full frequency table (15 rows)
#> ℹ Access $frequency_dist for per-value count distributions (86 rows)
#> ℹ Access $results for all samples and their horns indices (2,215 rows)
#> # Use `print(show = "all")` to see all elements
#> # Use `print(show = "none")` to hide elements

See closure_generate() for more details.

In addition to the bar plot, unsum offers an ECDF plot for CLOSURE results:

closure_plot_ecdf(data)

Empirical cumulative distribution function (ECDF) plot of `data`, the CLOSURE output. The curve rises most steeply at the first and last scale values, indicating a strongly polarized distribution.

Horns index variation

You may wonder about variability between the samples. Couldn’t there be some with a much lower or higher horns index than the overall mean horns? In this case, there would be a chance that the original data looked quite different from the average.

To check this, first use closure_plot_bar(). It shows the minimum and maximum possible variability as measured by the horns index, i.e., the average distributions of those samples with the lowest and highest horns index values:

closure_plot_bar(data, format = "percent")

Two barplots like the first, except they show the subsets of samples within `data` with minimum and maximum horns index values.

As you can see, the variability does not change very much. The distribution is starkly bimodal even with the lowest possible amount of variability.

In sum, the horns values are quite tightly confined. Wide variation among them seems to occur only if mean and sd have no decimal places.

Read and write

What if you have a huge object with CLOSURE results that you want to save? Write it to disk with closure_write():

# Using a temporary folder via `tempdir()` just for this example --
# you should use a real folder on your computer instead!
path_new_folder <- closure_write(data, path = tempdir())
#> 
#> ✔ All CLOSURE files written to:
#> /var/folders/z6/t79c19k13tl177xn70n2ybc80000gn/T//RtmpoNSybL/CLOSURE-3_5-1_8-80-1-5-up_or_down-5/

This stores the results using the highly efficient Parquet format. It will only take a tiny fraction of a CSV file’s disk space.

In your later session, read the data in from the folder to get the same CLOSURE list back:

data_new <- closure_read(path_new_folder)

A caveat: don’t modify the output of closure_generate() before passing it into other closure_*() functions. The latter need input with a very specific format, and if you manipulate the data between two closure_*() calls, these assumptions may no longer hold. Some checks are in place to detect alterations, but they may not catch all of them.