Reproducible compute sessions: pinning the full stack from R version to system libraries
A lockfile records which package versions ran. It notes the R version but can’t install it, and it never sees the system libraries or the OS underneath, any of which can quietly change your results.
A regulatory submission is filed. Eighteen months later, a query comes back from the FDA: reproduce the exact analysis behind a specific table. The analyst who ran it left the company a year ago.
The renv.lock file is still in the repository, faithfully listing every package version, and even the R version it ran on, 4.1. But 4.1 is no longer the default anywhere on the platform, and the lockfile only names it, it can’t reinstall it. The file also can’t tell you which BLAS implementation the system linked against, and that one was swapped out for a security patch months ago. The container the analysis actually ran in, if there ever was one, was never saved. The lockfile is accurate and useless at the same time.
This post is about making that scenario impossible by design, not by discipline.
What a lockfile doesn’t pin
renv (from Posit) for R, and pip’s requirements.txt or a lockfile like poetry.lock for Python, all do one job well: they record the exact package versions a project depends on. That’s necessary. It’s also only the top of the stack. Three things sit under the package list, and a lockfile pins none of them, even though any one can change your results.
- The R version. Base R’s own behavior shifts between releases. R 3.6.0 changed the default method behind sample(), so the same seeded code could return one result on 3.5 and another on 3.6, no error, not a line of the script touched. renv does record the R version in renv.lock, but as a label, not something it installs. The number tells you which R to find. It can’t rebuild it for you.
- The system libraries. These never show up in a DESCRIPTION file. R’s linear algebra runs through a BLAS library, and swapping one BLAS build for another can change the floating-point result of the same matrix operation (the two sum and parallelize in a different order). Anything touching spatial data, XML, or secure connections links against system libraries the same way. Patch the OS image, and a package built on top can shift without a single version number moving.
- Bioconductor. Less a layer of its own than a package ecosystem bolted to a specific R version. Try to upgrade R without moving to the matching Bioconductor release, and installs will either fail outright or quietly resolve to versions nobody validated together.
None of this is a knock on those tools. They solve the layer they were built for. The opening scenario happens because a correct, complete lockfile gets mistaken for a complete description of the environment.
The validated image is the unit of reproducibility
The fix is to stop treating the environment as everything except the package list. Treat it as a single artifact instead: one container image that pins the OS, the system libraries, the R version, and the Bioconductor release together, built once and referenced by digest. On a recent SCE build we made that image the unit of reproducibility, and everything downstream got simpler for it.
That image is built in layers. The base layer holds the operating system, the R runtime, the system libraries, and the packages that have cleared whatever risk-based validation your process asks for. That base gets tested, qualified, and locked. Because it’s referenced by digest, it can’t drift.
What actually gets validated shifts, too. You don’t re-validate the image every time it runs, because immutability already guarantees this run matches the last one. You validate the process that built the image: the definition reviewed, the packages checked, the pipeline controlled. Validate that once, and every image it produces inherits the result.
From there, the qualified image gets tagged to its validation state and pushed to a registry. It’s the same idea as a piece of calibrated equipment carrying a certificate, instead of being re-verified from scratch every time someone uses it. “Was this environment validated” becomes a lookup, not a reconstruction exercise. The tag is the human-readable validation label. The digest beneath it is the exact bytes a study pins.
One caveat worth planning for. A clinical program may need records retrievable fifteen years later, so the registry has to be treated as a retention system in its own right. Images stay resolvable by digest for as long as the study that points to them lives on, not garbage-collected on whatever schedule a registry defaults to.
Introducing a new R version without breaking old studies
You can’t just update the image in place. If “the current image” is a moving target, every study that depends on it inherits drift the moment someone rebuilds it. That’s the exact failure this whole approach exists to prevent.
The practical answer is parallel tracks, not one mutable image. A new R version or an updated system library becomes a new image, built and validated on its own, sitting alongside the ones already in use rather than replacing them. Existing studies keep pointing at the digest they were validated against. New work builds against the new image once it’s passed validation and been promoted from candidate to approved. A study’s environment changes only when someone deliberately re-points it, a recorded decision rather than a side effect of a rebuild.
The last piece is making sure analysts land on the right image without having to remember which one. That mapping, study to validated image, is project configuration, not tribal knowledge. It’s also where this connects to how people launch sessions day to day.
renv and Posit Package Manager together
Pinning the image handles the OS, the system libraries, the R version, and Bioconductor. The package layer still needs its own answer, and renv.lock alone doesn’t quite get there. A lockfile records what to install and which repository to fetch it from. The catch is that repository: point it at the live CRAN mirror and the old versions simply aren’t there anymore, especially for older R releases.
That’s the gap Posit Package Manager’s dated snapshots close. Point renv at a specific calendar-date snapshot instead of the live CRAN mirror, and the lockfile becomes self-sufficient. Rebuilding the package environment years later just means reading the lockfile, which already points at a snapshot that isn’t going anywhere. renv.lock plus a dated snapshot is what turns “we recorded the versions” into “we can rebuild the exact environment from this one file.”
Three artifacts, one reproducible result
Put the image and the package layer together, and a whole analysis run reduces to three references:
- the container image digest, for the environment,
- the renv.lock resolved against its Package Manager snapshot, for the packages,
- the git commit of the analysis code, for the code.
Recorded once, when the analysis runs, these three are what “reproducible” means in practice.
So when someone asks you to prove that a submission’s analysis used the same code, the same packages, and the same environment as the original run, the answer isn’t a reconstruction project. Pull the image by digest, restore from the lockfile, check out the commit, run. All three still resolve, because none of them depends on a live upstream staying put. The FDA query from the opening scenario stops being a scramble and becomes three lookups.
Where Posit Workbench fits
None of this holds if analysts can still launch a session from whatever image happens to be lying around. Posit Workbench’s session launcher closes that gap. Administrators define the set of validated images available as launch options, each one a specific R version, Bioconductor release, and set of system libraries, already qualified. A session starts from the image its project is tied to, not from one an analyst picked because it happened to have a package they wanted.
The choice of editor rides on top of that. RStudio, VS Code, and JupyterLab are layers over the same validated base, three windows onto one compute engine. Which one an analyst opens, or whether it gets updated next quarter, doesn’t touch the runtime, the packages, or the system libraries underneath. So it doesn’t put the study’s reproducibility in question.
That’s what makes the validated image more than a registry entry nobody looks at. It’s the environment a statistician’s session actually boots into, every time, for that study, until someone makes the deliberate, recorded decision to move it forward. The image lifecycle and the day-to-day session become the same mechanism, not two systems kept in sync by hand.
Where this takes you
With the full stack pinned, that eighteen-month-later reproduction stops being a question of who still remembers what was installed. It’s a container digest, a lockfile, and a commit hash, all sitting in the study record before anyone starts asking. That’s the difference between a platform that can answer a regulator’s question and one that has to go looking for the answer.
This post expands one part of Appsilon’s guide, The AI-Ready SCE: A Pharma Guide to Building Modern Environments, which lays out the full platform from validation through AI.

