Reproducibility and sharing

DATA → MODEL → ESTIMATE ⇄ EVALUATE → SIMULATE → REPORT

Where you are

The last step of the workflow makes the analysis durable. Results should be saved in a form colleagues and tools can open. Every number in the report should be traceable to a model file, a dataset and a software version. Re-running the analysis should give the same answer. This chapter covers ferx’s portable fit files and the provenance a fit records.

The data

cov <- ferx_example("two_cpt_oral_cov")
run_dir <- book_tempdir("runs")

Minimal runnable call: saving a fit

ferx_fit() can save the fit while it runs: output is the file path, and include_data = TRUE also embeds the dataset. A path without an extension gets .fitrx:

fit <- ferx_fit(cov$model, cov$data, output = file.path(run_dir, "run1"), verbose = FALSE)
list.files(run_dir)
#> [1] "run1.fitrx"

ferx_save_fit() does the same for an existing fit, with the arguments fit, output and include_data. It returns the fit invisibly, so it can sit in a pipe:

ferx_save_fit(fit, output = file.path(run_dir, "run1_with_data"), include_data = TRUE)

A .fitrx file is a zip archive of JSON and CSV files, readable from any language with zip, JSON and CSV support:

unzip(file.path(run_dir, "run1_with_data.fitrx"), list = TRUE)[, c("Name", "Length")]
#>              Name Length
#> 1   manifest.json    262
#> 2        fit.json   9245
#> 3        ebes.csv   3572
#> 4 predictions.csv  32828
#> 5      covtab.csv   6316
#> 6      model.ferx   1107
#> 7    warnings.txt    124
#> 8        data.csv  10680

It contains the model source (model.ferx), the estimates and run information (fit.json), per-subject random effects, per-observation predictions, the covariate table, the warnings, and with include_data = TRUE the dataset. The ferx-core .fitrx format page documents the schema.

Reading the result: loading a fit

ferx_load_fit() reads a .fitrx file back as a ferx_fit object. Its path may omit the extension:

loaded <- ferx_load_fit(file.path(run_dir, "run1"))
class(loaded)
#> [1] "ferx_fit"
all.equal(fit$estimates, loaded$estimates)
#> [1] TRUE
all.equal(fit$sdtab, loaded$sdtab)
#> [1] TRUE

The loaded fit can be used like the original. It can be simulated from, or given a SIR step (here the model and data files are unchanged, so the integrity check passes):

sim <- ferx_simulate(cov$model, cov$data, n_sim = 2, seed = 1, fit = loaded)
nrow(sim)
#> [1] 600
ferx_sir(loaded, sir_seed = 1)$sir_ess
#> [1] 17.86333

With include_data = TRUE, the loaded fit points to a copy of the embedded dataset, so the bundle is self-contained:

loaded_data <- ferx_load_fit(file.path(run_dir, "run1_with_data.fitrx"))
file.exists(loaded_data$data_path)
#> [1] TRUE

What a .fitrx file does not keep

The file stores the fields of the cross-language schema, not every element of the R fit object:

setdiff(names(fit), names(loaded))
#>  [1] "warnings_severity"      "warnings_category"      "warnings_message"      
#>  [4] "warnings_source_method" "sir_resamples"          "sir_resamples_n"       
#>  [7] "sir_resamples_dim"      "importance_sampling"    "impmap_trace"          
#> [10] "bayes"                  "cond_dist"              "individual_estimates"  
#> [13] "saem_n_subjects_hmc"    "neural_networks"        "obs_time_range"        
#> [16] "final_gradient"         "optimizer_label"        "n_starts"              
#> [19] "multi_start_seed"       "saem_seed"              "sir_seed_used"         
#> [22] "imp_seed"               "bloq_method_label"      "outer_maxiter"         
#> [25] "outer_gtol"             "inits_from_nca"         "exclusions"            
#> [28] "eigenvalues"            "condition_number"       "eta_normality"         
#> [31] "warnings_structured"

Warning categories are an example. After loading, the warnings are still there, but their category is general:

ferx_get_warnings(fit, as_df = TRUE)$category
#> [1] "dw_autocorrelation"
ferx_get_warnings(loaded, as_df = TRUE)$category
#> [1] "general"

Keep what you need from the missing fields (for example individual_estimates or condition_number) in your own tables. Or save the R object as well, with saveRDS().

Recording provenance

A fit records where it came from:

data.frame(
  item = c("model file", "data file", "model SHA-256", "data SHA-256", "engine version", "ferx-r version"),
  value = c(basename(fit$model_path), basename(fit$data_path), substr(fit$model_hash, 1, 16),
            substr(fit$data_hash, 1, 16), fit$ferx_version, as.character(packageVersion("ferx")))
)
#>             item                 value
#> 1     model file two_cpt_oral_cov.ferx
#> 2      data file  two_cpt_oral_cov.csv
#> 3  model SHA-256      108876c28cf58c52
#> 4   data SHA-256      d7c03b6285d9c6a5
#> 5 engine version                 0.4.0
#> 6 ferx-r version                 0.3.0

fit$model_text holds the model source. The hashes are what ferx_sir() and ferx_covariance() check before reusing a fit: if the model or data file has changed, they stop.

Package versions alone do not identify a development build. This book records the ferx-r commit it was built with (6e2f701, see Installing ferx). Do the same in your reports, together with the R session information:

sessionInfo()$running
#> [1] "Ubuntu 24.04.5 LTS"
c(R = R.version.string, ferx = as.character(packageVersion("ferx")))
#>                              R                           ferx 
#> "R version 4.4.3 (2025-02-28)"                        "0.3.0"

A run log (Estimation methods and controlling the fit) is a readable record to keep next to the .fitrx file:

writeLines(ferx_runlog(fit, verbose = FALSE), file.path(run_dir, "run1_runlog.txt"))
list.files(run_dir)
#> [1] "run1_runlog.txt"      "run1_with_data.fitrx" "run1.fitrx"

Options that matter

Function Argument Meaning
ferx_fit() output Save the fit to this .fitrx path when fitting
include_data Embed the dataset in the file (only with output)
ferx_save_fit() fit The fit to save
output File path (.fitrx appended when there is no extension)
include_data Embed the dataset
ferx_load_fit() path .fitrx file, with or without extension

Every stochastic step in the book takes a seed. Set it explicitly so results are reproducible:

Where Seed
ferx_simulate(), ferx_simulate_with_uncertainty(), ferx_simulate_adaptive(), ferx_calc_npde() seed argument
ferx_bootstrap() seed argument (results independent of threads)
ferx_sir() sir_seed argument
SAEM, IMP, IMPMAP, Bayes, multi-start fits seed, imp_seed, impmap_seed, bayes_seed, multi_start_seed settings (Estimation methods and controlling the fit)

Variants

A project layout

ferx does not require a particular layout. One that keeps runs traceable:

project/
├── data/            # analysis datasets, as used by the fits
├── models/          # run1.ferx, run2.ferx, ... (versioned model files)
├── runs/            # run1.fitrx, run1_runlog.txt, search and bootstrap directories
├── scripts/         # R scripts or Quarto documents that run the analysis
└── report/          # tables, figures and the report source

Keep search directories (Model development and selection) and bootstrap directories (Parameter uncertainty) under runs/: they contain every candidate model and can be resumed. A Quarto document that fits, evaluates and reports, like the chapters of this book, is itself a reproducible record of the analysis.

Pitfalls

  • Reloaded fits are not complete R fit objects. Check that the fields you need survive ferx_save_fit()/ferx_load_fit(), and do not rely on warning categories after loading.
  • Changing a model or data file invalidates follow-up steps. ferx_sir() and ferx_covariance() stop when the hashes no longer match. Version model files (run2.ferx) instead of editing them in place.
  • Record the build. Output can change between ferx builds. Record the ferx-r commit or release with every run.

Summary

  • Save fits with ferx_fit(output = ...) or ferx_save_fit(), load them with ferx_load_fit(). Use include_data = TRUE for self-contained files.
  • A fit records its model text, file hashes and engine version. Add the ferx-r commit and session information to reports.
  • Set seeds for every stochastic step, and keep model files, run files, logs and search directories together.

This completes The analysis workflow. The scenario Parts that follow apply it to specific modeling problems, starting with Structural models: analytical and ODE.

TipReference
  • R help: ?ferx_save_fit, ?ferx_load_fit, ?ferx_fit (output, include_data)
  • ferx-core: .fitrx file format