CCTMB #003 Sizing a Clinical Trial: Can't Test? – Don't Test
- Leila Janani

- 6 days ago
- 5 min read
The design stage of a clinical trial is my favourite part of being a medical statistician, as there is no single perfect design. Instead, it offers a chance to discover and explore differing options. One aspect that many trialists struggle with is how to size the trial. This is unsurprising, as while the basic sample size calculation is simple and quick to perform, there is a lot more to take into consideration. It all starts with precisely defining a research question, and this will guide the estimates for some of the sample size parameters that can be plugged into a sample size calculator.
This blog is about the common phenomenon that occurs when researchers don’t like the answer of their first sample calculation; when they feel it is too large. They go back and change the sample size parameters until the calculation results in a number that is more palatable to them. This shimmy has been called the Sample Size Tango. It’s this issue that I want to focus on here, because while it’s a well-recognised phenomenon, the consequence that this practice has on wasted resources and knock-on impact for downstream misinterpretation of trial results does not seem to be well understood.
This practice is especially common in academic trials. Why? Because investigators are often working within tight budget requirements set by funders and constraints for needing a timely answer, making it genuinely difficult to recruit the number of participants needed for an adequately powered late-phase frequentist trial.
So, what’s the problem here? Surely any data is better than no data? Yes - I agree, but NOT when we are hypothesis testing. You see, the sample size parameter that trialists most often alter to reduce their sample size is the ‘minimum clinically important difference’ (MCID) that they will be able to detect. They choose this parameter as it has the biggest impact on reducing sample size, but this means that they then embark on an underpowered trial. Consequently, if the true intervention effect is smaller than anticipated but still clinically important, the trial is highly likely to be unable to detect it.

The logic just does not add up. We are setting our own exam paper in frequentist trials. So why set one we know we can’t pass for an important effect? The knock-on effect is to see these trials get incorrectly and repeatedly summarised as ‘there was no difference’ and any nuance for interpretation lost.
The analogy is the same for non-inferiority trials, but arguably worse, as non-inferiority trials generally require larger sample sizes to rule out smaller effects, and the consequence of the interpretation is more harmful, as we can easily conclude an intervention is not inferior when it can be harmful.
So why do we allow this repeated bad design practice to occur? I would suggest that it’s because we are all sympathetic to the cause. We know there is a practical limit on the number of participants, centres, countries and funding we have access to, and we need the results within a meaningful period of time. So, what should we do instead? Well, there are options available, but I would strongly argue that setting ourselves up for a fall with unrealistic hypothesis tests to pass is not one of them. We should also call this practice out when we see it at the design and funding stage (once past this stage, it’s too late!).
One aspect that we can often make improvements to is statistical. Get some statistical advice on this. For example, we could:
use information rich outcomes, e.g. using continuous or ordinal outcome measures and not dichotomising lovely continuous scales.
collect repeated measurements of the outcome to add more information.
decide to test at a higher significance level, and while that means a higher chance for concluding the intervention works when it does not, that may be a suitable trade-off in some circumstances (for example, in Phase II trials).
consider whether the comparator (control) group is the right one (for example, a smaller sample size is needed when comparing to a placebo versus another active intervention).

While all of these suggestions will reduce the sample size, they will also have an impact on the research question being answered. Therefore, it will be important to circle back and update the research question in light of the design and the sample size calculation.
Adaptive trial designs are also an option. They may not reduce the sample size for an individual trial, but they will have gains overall as they reduce the expected sample size.
One of my favourite papers exploring the topic of small sample size is by Parmar et al. 2016 as it provides a practical framework for addressing ‘How do you design randomised trials for smaller populations?’ . It considers the reason for not wanting to go ahead with the larger sample size and explores opportunities to maximise recruitment before examining the trade-offs between the research question, operational considerations and trial design when considering a reduction in the sample size. However, even with all these steps, the sample size may still be too large to practically recruit in a feasible time and acceptable funding envelope. In these situations, one compelling alternative is to adopt a Bayesian framework.
A Bayesian approach has the philosophy that evidence should accumulate. It also allows us to combine evidence generated by the trial with relevant existing knowledge, rather than treating each trial as though it exists in isolation.
In the pure Bayesian framework, we don’t need to set up artificial hurdles over which to jump based on a pre-specified hypothesis but view the evidence as it is and use the posterior distribution of the treatment effect to summarise the evidence and use it to inform our judgement and decision-making. This shifts our attention to the distribution of the treatment effect rather than evaluating the observed data against what would be expected under the assumption of no treatment effect, as we do in the Frequentist framework. We can use the posterior distribution to answer clinically meaningful questions directly and use this more flexibly, such as what is the probability that a treatment provides a worthwhile benefit.
As a result of the obvious advantages in our field, I find myself more committed to using a Bayesian framework. The underlying philosophy feels natural and suited for our purpose of evaluating interventions. For many years, uptake was slow as the implementation of theory and practice in the past has been complex. Today, there are many tools, training and software to make it easier to apply Bayesian methods in practice.
To stop the harm we are doing by missing effective interventions and labelling ineffective interventions as effective, we need to stop dancing the Sample Size Tango. Instead, look at statistical efficiency in the design and the research question you can answer, and consider the most suitable statistical framework to use.
My advice in short… Can't Test? – Don't Test
References and resources
Cook JA, Julious SA, Sones W. et al. DELTA2 guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial. BMJ. 2018 Nov 5; doi: 10.1136/bmj.k3750.
Parmar MK, Sydes MR, Morris TP. How do you design randomised trials for smaller populations? A framework. BMC Med. 2016 Nov 25;14(1):183. doi: 10.1186/s12916-016-0722-3.
Berry, D.A. (1993), A case for bayesianism in clinical trials. Statist. Med., 12: 1377-1393. https://doi.org/10.1002/sim.4780121504



Comments