DATA → MODEL → ESTIMATE ⇄ EVALUATE → SIMULATE → REPORT
Reproducibility and sharing
Where you are
The last step of the workflow makes the analysis durable. Results should be saved in a form colleagues and tools can open. Every number in the report should be traceable to a model file, a dataset and a software version. Re-running the analysis should give the same answer. This chapter covers ferx’s portable fit files and the provenance a fit records.
The data
cov <- ferx_example("two_cpt_oral_cov")
run_dir <- book_tempdir("runs")Minimal runnable call: saving a fit
ferx_fit() can save the fit while it runs: output is the file path, and include_data = TRUE also embeds the dataset. A path without an extension gets .fitrx:
fit <- ferx_fit(cov$model, cov$data, output = file.path(run_dir, "run1"), verbose = FALSE)
list.files(run_dir)
#> [1] "run1.fitrx"ferx_save_fit() does the same for an existing fit, with the arguments fit, output and include_data. It returns the fit invisibly, so it can sit in a pipe:
ferx_save_fit(fit, output = file.path(run_dir, "run1_with_data"), include_data = TRUE)A .fitrx file is a zip archive of JSON and CSV files, readable from any language with zip, JSON and CSV support:
It contains the model source (model.ferx), the estimates and run information (fit.json), per-subject random effects, per-observation predictions, the covariate table, the warnings, and with include_data = TRUE the dataset. The ferx-core .fitrx format page documents the schema.
Reading the result: loading a fit
ferx_load_fit() reads a .fitrx file back as a ferx_fit object. Its path may omit the extension:
The loaded fit can be used like the original. It can be simulated from, or given a SIR step (here the model and data files are unchanged, so the integrity check passes):
sim <- ferx_simulate(cov$model, cov$data, n_sim = 2, seed = 1, fit = loaded)
nrow(sim)
#> [1] 600
ferx_sir(loaded, sir_seed = 1)$sir_ess
#> [1] 17.86333With include_data = TRUE, the loaded fit points to a copy of the embedded dataset, so the bundle is self-contained:
loaded_data <- ferx_load_fit(file.path(run_dir, "run1_with_data.fitrx"))
file.exists(loaded_data$data_path)
#> [1] TRUEWhat a .fitrx file does not keep
The file stores the fields of the cross-language schema, not every element of the R fit object:
setdiff(names(fit), names(loaded))
#> [1] "warnings_severity" "warnings_category" "warnings_message"
#> [4] "warnings_source_method" "sir_resamples" "sir_resamples_n"
#> [7] "sir_resamples_dim" "importance_sampling" "impmap_trace"
#> [10] "bayes" "cond_dist" "individual_estimates"
#> [13] "saem_n_subjects_hmc" "neural_networks" "obs_time_range"
#> [16] "final_gradient" "optimizer_label" "n_starts"
#> [19] "multi_start_seed" "saem_seed" "sir_seed_used"
#> [22] "imp_seed" "bloq_method_label" "outer_maxiter"
#> [25] "outer_gtol" "inits_from_nca" "exclusions"
#> [28] "eigenvalues" "condition_number" "eta_normality"
#> [31] "warnings_structured"Warning categories are an example. After loading, the warnings are still there, but their category is general:
ferx_get_warnings(fit, as_df = TRUE)$category
#> [1] "dw_autocorrelation"
ferx_get_warnings(loaded, as_df = TRUE)$category
#> [1] "general"Keep what you need from the missing fields (for example individual_estimates or condition_number) in your own tables. Or save the R object as well, with saveRDS().
Recording provenance
A fit records where it came from:
data.frame(
item = c("model file", "data file", "model SHA-256", "data SHA-256", "engine version", "ferx-r version"),
value = c(basename(fit$model_path), basename(fit$data_path), substr(fit$model_hash, 1, 16),
substr(fit$data_hash, 1, 16), fit$ferx_version, as.character(packageVersion("ferx")))
)
#> item value
#> 1 model file two_cpt_oral_cov.ferx
#> 2 data file two_cpt_oral_cov.csv
#> 3 model SHA-256 108876c28cf58c52
#> 4 data SHA-256 d7c03b6285d9c6a5
#> 5 engine version 0.4.0
#> 6 ferx-r version 0.3.0fit$model_text holds the model source. The hashes are what ferx_sir() and ferx_covariance() check before reusing a fit: if the model or data file has changed, they stop.
Package versions alone do not identify a development build. This book records the ferx-r commit it was built with (6e2f701, see Installing ferx). Do the same in your reports, together with the R session information:
sessionInfo()$running
#> [1] "Ubuntu 24.04.5 LTS"
c(R = R.version.string, ferx = as.character(packageVersion("ferx")))
#> R ferx
#> "R version 4.4.3 (2025-02-28)" "0.3.0"A run log (Estimation methods and controlling the fit) is a readable record to keep next to the .fitrx file:
writeLines(ferx_runlog(fit, verbose = FALSE), file.path(run_dir, "run1_runlog.txt"))
list.files(run_dir)
#> [1] "run1_runlog.txt" "run1_with_data.fitrx" "run1.fitrx"Options that matter
| Function | Argument | Meaning |
|---|---|---|
ferx_fit() |
output |
Save the fit to this .fitrx path when fitting |
include_data |
Embed the dataset in the file (only with output) |
|
ferx_save_fit() |
fit |
The fit to save |
output |
File path (.fitrx appended when there is no extension) |
|
include_data |
Embed the dataset | |
ferx_load_fit() |
path |
.fitrx file, with or without extension |
Every stochastic step in the book takes a seed. Set it explicitly so results are reproducible:
| Where | Seed |
|---|---|
ferx_simulate(), ferx_simulate_with_uncertainty(), ferx_simulate_adaptive(), ferx_calc_npde()
|
seed argument |
ferx_bootstrap() |
seed argument (results independent of threads) |
ferx_sir() |
sir_seed argument |
| SAEM, IMP, IMPMAP, Bayes, multi-start fits |
seed, imp_seed, impmap_seed, bayes_seed, multi_start_seed settings (Estimation methods and controlling the fit) |
Variants
A project layout
ferx does not require a particular layout. One that keeps runs traceable:
project/
├── data/ # analysis datasets, as used by the fits
├── models/ # run1.ferx, run2.ferx, ... (versioned model files)
├── runs/ # run1.fitrx, run1_runlog.txt, search and bootstrap directories
├── scripts/ # R scripts or Quarto documents that run the analysis
└── report/ # tables, figures and the report source
Keep search directories (Model development and selection) and bootstrap directories (Parameter uncertainty) under runs/: they contain every candidate model and can be resumed. A Quarto document that fits, evaluates and reports, like the chapters of this book, is itself a reproducible record of the analysis.
Pitfalls
-
Reloaded fits are not complete R fit objects. Check that the fields you need survive
ferx_save_fit()/ferx_load_fit(), and do not rely on warning categories after loading. -
Changing a model or data file invalidates follow-up steps.
ferx_sir()andferx_covariance()stop when the hashes no longer match. Version model files (run2.ferx) instead of editing them in place. - Record the build. Output can change between ferx builds. Record the ferx-r commit or release with every run.
Summary
- Save fits with
ferx_fit(output = ...)orferx_save_fit(), load them withferx_load_fit(). Useinclude_data = TRUEfor self-contained files. - A fit records its model text, file hashes and engine version. Add the ferx-r commit and session information to reports.
- Set seeds for every stochastic step, and keep model files, run files, logs and search directories together.
This completes The analysis workflow. The scenario Parts that follow apply it to specific modeling problems, starting with Structural models: analytical and ODE.
TipReference
- R help:
?ferx_save_fit,?ferx_load_fit,?ferx_fit(output,include_data) - ferx-core: .fitrx file format