Case Study

A GxP-Ready Statistical Computing Environment for Regulatory Submissions

A mid-sized pharmaceutical company replaced an ad hoc analytics setup with a production-grade statistical computing environment supporting validated work and clinical analyses intended for FDA submission. Approximately 250 existing workflows were re-architected into governed processes, from vendor data intake through regulatory archival.

Client:
Mid-sized pharmaceutical company (clinical-stage)
Services:
Technologies used:
R
Python
Terraform
Posit Connect
Posit Workbench
Posit Package Manager
astellas
Genmab
merck
johnson and johnson
World Health Organisation
Kenvue
Phuse
Phuse
Phuse
Phuse
Phuse
astellas
Genmab
merck
johnson and johnson
World Health Organisation
Kenvue
Phuse
Phuse
Phuse
Phuse
Phuse

Table of contents

Before:
The biometrics team could not use their existing SCE for clinical trial analyses efficiently and lacked a practical way to do their jobs. The old SCE had been designed around IT infrastructure and QA requirements without adequately supporting the biometrics team’s daily work. About 250 workflows remained loosely documented.
After:
A production-grade, GxP-ready platform where the biometrics team runs validated analyses for FDA submissions. About 250 workflows consolidated into one governed lifecycle from vendor data intake to regulatory archival. Exploratory work and regulatory outputs on separate paths, each with the right level of control.

A mid-sized pharmaceutical company replaced an ad hoc analytics setup with a production-grade statistical computing environment supporting validated work and clinical analyses intended for FDA submission. Approximately 250 existing workflows were re-architected into governed processes, from vendor data intake through regulatory archival.

At a Glance

  • A production-grade, GxP-ready SCE supporting validated clinical analysis.
  • Approximately 250 existing workflows re-architected into governed processes.
  • Separate exploratory and regulatory workflows, with validation effort matched to risk

About the Project

Our client's biometrics and data science team analyzes clinical trial data that ends up in FDA submissions. The infrastructure behind that work did not match its stakes. Analyses ran on shared servers, data moved between people and systems by hand, and an audit trail existed only if someone assembled it after the fact. 

Working with Appsilon, the client now has a validated statistical computing environment built on an open-source stack. Submission analyses run under GxP controls with a complete audit trail. About 250 existing workflows were re-architected into a governed set of processes covering the full data lifecycle, and validation effort is concentrated on the components that carry regulatory risk.

The Problem: Clinical Analysis Depended on Shared Servers and Manual Handoffs

The company's biometrics and data science teams needed to analyze clinical trial data intended for FDA submission. Their existing setup relied on shared servers and manual file transfers, without a built-in audit trail.

That left a gap between the work scientists needed to do and the controls required to support regulatory use. The company needed a secure place to receive clinical data, run analyses, publish analysis applications, and retain the resulting records with traceability throughout.

A previous implementation attempt had not delivered a usable solution. Data science, quality, regulatory, and IT stakeholders also had different expectations of what the replacement needed to provide. Success meant delivering an environment scientists could use while addressing the requirements of the teams responsible for its oversight.

Clinical data destined for a regulator has to be traceable, reproducible, and produced under controlled access. Shared servers and manual handoffs setup was creating several issues:

  1. Traceability depended on people, not systems: who changed what, when, and why has to be reconstructed later. That is the first question an inspector asks.
  2. The team had a platform in place, but could not use it to do the analyses they needed since the environment in place was set up around IT and validation requirements without adequately addressing how biometrics works with clinical trial data. 
  3. Every analysis carried the same compliance burden, and exploratory work runs at the speed of validated work, so scientists either had to slow down or work outside the controlled environment.

The Solution: an SCE Designed around Clinical Work and Regulatory Requirements

Working with Appsilon, the company built a statistical computing environment (SCE): a shared platform for receiving clinical data, running analyses, publishing analysis applications, and retaining records. Access controls and auditable processes governed how teams worked with data throughout its lifecycle.

The implementation combined knowledge of clinical data workflows and GxP validation with software engineering. Project-scoped computing resources and access controls governed who could work with data. Automated pipelines recorded an auditable history, and scientists could publish analysis applications themselves within defined governance controls. These capabilities addressed the needs of scientists alongside the oversight requirements of quality, regulatory, and IT teams.

A prototype delivered in six weeks allowed the client's team to assess the approach in a working environment. It provided a basis for the subsequent enterprise rebuild.

Step 1: A Working Environment Under Real Controls

Six weeks from project start

Instead of a long discovery phase followed by a year-long validation cycle, the first six weeks produced an environment the team could use for real work:

  1. One home base for applications and data
  2. Project-scoped compute with access controls set per study
  3. Automated data pipelines with a complete audit trail
  4. Governed self-service publishing for analysis apps

This settled the question the previous attempt had left open and gave all four stakeholder groups a concrete system to align on, rather than a specification.

Step 2: The Greenfield Rebuild

Enterprise-scale platform

With the concept proven, the platform was rebuilt as a scalable enterprise SCE. The design decisions here are what make it work for regulated analyses:

  • The enterprise rebuild addressed approximately 250 existing workflows, many loosely documented. These were re-architected into governed processes covering vendor data intake, analysis, and regulatory archival.
  • Risk-based validation directed effort according to regulatory risk. The platform also separated exploratory analysis from formal, validated regulatory outputs. Scientists could investigate data without applying the full validation process to every experiment, while work intended for regulatory use followed the appropriate controls.
  • Exploration separated from regulatory outputs. Scientists explore freely; only outputs headed for a submission go through the formal validated path. Both live on the same platform with the same audit trail.
  • Open-source stack, with no vendor lock-in, ensuring lower licensing cost, and a large community maintaining every layer. The R packages the platform depends on were validated with Axon.R, Appsilon's GxP-aligned R package validation framework, that produces traceable evidence for QA review and supports controlled package approval.
  • AI-ready architecture that gave the company a basis for extending its analytical capabilities as its needs developed, without re-architecting the whole platform.
The governed data lifecycle. Exploratory and validated work share one platform and one audit trail; only validated outputs proceed to regulatory archival.

Business Impact: A Working Platform for Validated Clinical Analysis

Support for Submission Work

The company established a production-grade, GxP-ready SCE supporting its biometrics and data science teams in preparing analyses intended for FDA submission. Governance extended across the clinical data lifecycle. Biometrics Teams can run clinical analyses with the checks and records needed to support regulatory work, all within one shared environment.

Clear Record of What Happened to the Data

Automated processes record how data moves through the platform. Teams can follow its history from arrival through analysis and storage.

Less Validation Work for Exploratory Analysis

Exploratory work is separated from regulatory outputs, allowing Scientists to explore data without putting every experiment through the full validation process. Analyses intended for regulatory use follow stricter checks, with validation effort focused on higher-risk work.

Future Development 

Open-source software reduces reliance on a single platform vendor. The company can also add AI and machine learning capabilities as its needs grow.

Explore statistical computing environments for pharma and biotech.

Explore Possibilities

Share Your Data Goals with Us

From advanced analytics to platform development and pharma consulting, we craft solutions tailored to your needs.