Six years later, we ran the experiment again 

In 2020, the first thing we did as a company was replicate the CDISC pilot in R. The question at the time was whether R could hold up for clinical reporting at all. Instead of guessing, we checked. We took the pilot’s statistical appendix — all 30 safety and efficacy tables (a reference most people in the industry already know) — and rebuilt it with tidyverse and {huxtable}. Out of the box, not all the tooling was there, so we built and released {pharmaRTF}. Six years on, tooling has evolved beyond recognition. We ran the experiment again, and this time we handed all 30 tables to Claude Code. 

The rules mattered more than the tools 

For the rebuild, we wanted to compare against the original outputs, but without bias against the existing code. The 2020 programs were sitting right there, and they served strictly as reference material for methodology. The target was the rendered table, not the program that produced it. For the tech stack, we used {tplyr2} for the summaries, {clinify} for the outputs, {mmrm} for the repeated measures analysis, and {coin} for the Cochran–Mantel–Haenszel tests. 

The process we designed and followed did most of the work. The instructions were specific: use {tplyr2} wherever a layer could carry the summary, and use {clinify} for every output. An agent left to its own devices will write bespoke code that works while teaching you nothing. 

We built the method for checking results up front, with correctness set up to be tested twice through numerical and visual QC. Numbers were extracted from every table and compared cell by cell. Then both documents were rendered to PDF through the same engine and compared visually. Single-page tables got a masked pixel diff. Multipage tables got a pagination-agnostic content check that compares unique body lines and controls for line wraps across pages, so re-pagination cannot hide a changed value. The visual comparison returns a percentage difference rather than a verdict. Interpreting that number is a judgment call. The goal of two-step QC was to verify if each table computes correctly, and that it reads correctly to the person who opens it. 

All 30 tables were rebuilt and verified. 

The process also found defects in the 2020 outputs. Table 14-6.06 carried a CMH p-value of 0.015 that came from a count vector passed into the test in the wrong shape. The correct value is 0.405, confirmed in two independent ways. Where a modern method produces a different number, the difference is documented rather than cloned. 

Develop, check, repeat 

The rebuild was highly iterative, with the goal of replacing bespoke pilot code with reusable package functionality wherever reasonable. It turned into a powerful feedback loop, because every handwritten workaround pointed to a missing feature in the package underneath it. Each bug became an issue with a runnable reproducible example (known as a “reprex” in the R community), followed by a pull request and a release. The fixes shipped in {tplyr2} 0.2.0 and {clinify} 0.4.0 were adopted back into the pilot programs and re-verified to keep every output identical. All but three, which are structurally bespoke tables, now build directly from the packages with very little custom code.  

The cycle was the same each time. Finish the tables and get them byte-identical. Round up everything that needs a workaround, such as bugs or enhancement opportunities in {tplyr2} or {clinify}, and file them as issues on GitHub. Pull those issues down into separate Claude Code sessions for each package, work through them, and submit the pull requests. Then pull the released changes back into the pilot, have the agent re-evaluate every table against the new capabilities, and either accept what came back or file the next round of comments and bugs. Repeat. 

The study work improved the infrastructure underneath it, and the infrastructure then improved the study work. As an industry, we tend to treat study programming as disposable. Write it, validate it, archive it, and never look at it again. This experiment showed that 30 tables of study work also shipped two significant package releases. 

What it cost 

One subscription, around $250 a month. We never hit the plan limits, even while rebuilding 30 tables and fixing and releasing two packages on a standard individual plan. That is worth highlighting, because stories of runaway cost dominate the conversation about agentic work.  

This cost was also not optimized, so there is more room for even more efficient outcomes. The agent pre-built tooling that rendered, diffed, and reported, and it worked inside that. Here the approach was to build the harness properly, scope it, and separate the review step from the generation step. The cost profile improves further, which is the real lesson on cost. The expensive part of agentic work is not solely tokens; it’s also the absence of deterministic tooling around the agent. 

So, would we still write everything twice? 

Let’s look at what the process actually did. It checked numeric accuracy independently of the original code. It checked visual correctness, which double programming alone doesn’t do, and which typically waits for senior review. It documented every deliberate difference instead of silently cloning a defect. And it flagged a wrong p-value in work that had already passed the review. 

Double programming was the best available mechanism when the only way to get a second opinion was a second person. That is no longer the only way. To be clear, agents shouldn’t replace QC. Independence is the requirement, and writing the same program twice is one implementation of it — one we chose under constraints that no longer hold. So how can tooling like this improve our processes today? 

The outputs, programs, and documentation from this rebuild are public: atorus-research.github.io/CDISC_pilot_replication 

Want to explore what this could look like for your team’s study programming? 

FAQs

Yes, with the right process around it. In 2026, Atorus used Claude Code to rebuild all 30 safety and efficacy tables from the CDISC pilot in R. Every table was verified cell by cell against the original outputs and checked visually. Accuracy came from the rules and the QC harness, not the agent alone. Specific package instructions and independent checks stopped the agent from writing bespoke code nobody could reuse.

Not as a shortcut, but it changes how independence is achieved. Double programming is the best way to get a second opinion when the only option is a second programmer. Independence is the actual requirement. In the Atorus rebuild, agent-driven QC checked numbers independently of the original code and also checked visual correctness.

Test correctness twice: numerically and visually. Atorus compared every table cell by cell against the original outputs, then rendered both versions to PDF and compared them visually, using a pixel diff for single-page tables and a pagination-agnostic content check for multipage tables.

One standard individual subscription. It covered all 30 safety and efficacy tables plus fixing and releasing two R packages, and the team never hit plan limits. The cost was not optimized.

The Atorus CDISC pilot rebuild used {tplyr2} for summaries, {clinify} for outputs, {mmrm} for repeated measures analysis, and {coin} for Cochran–Mantel–Haenszel tests. Gaps found during the rebuild shipped as fixes in {tplyr2} 0.2.0 and {clinify} 0.4.0. Outputs, programs, and documentation are public at atorus-research.github.io/CDISC_pilot_replication.

Mike Stackhouse
Chief Innovation Officer, Atorus

Michael Stackhouse leads innovation in clinical data engineering and analytics, specializing in bridging modern data science frameworks with compliant, scalable solutions for the life sciences industry.  With extensive expertise in CDISC and open-source technologies, his work spans automation, data science, big data solutions, and clinical analytics within regulated environments.  An award-winning industry leader, Michael is a frequent invited speaker, panelist, and published contributor in advancing the future of clinical data science. 

Aga

Aga Rasinska
Director of Strategy, Atorus

With more than a decade of hands-on experience in bioinformatics, data science, and program leadership, Aga Rasinska brings a dual perspective that bridges business strategy and technical innovation, transforming complex clinical and omics data challenges into scalable, results-driven solutions. Her expertise spans project and product management, change leadership, and data-driven decision-making within regulated life sciences environments. She is also a frequent industry speaker, panelist, and contributor focused on the intersection of science, data, and business strategy. 

Back to Blog