CCTMB #004 Rethinking Evidence in Critical Care Trials: Perspectives from the Critical Care Reviews Meeting
- Leila Janani

- 19 hours ago
- 7 min read
By Manel Benlloch Guajardo and Martin Orr, Clinical Trials Statisticians at Imperial Clinical Trials Unit (ICTU)
In June, members of the ICTU team attended the Critical Care Reviews meeting in Belfast. This stunning conference mixed groundbreaking critical care trials and independent editorials with the chance to experience first-hand the feeling of being in the Titanic grand staircase. A true blend of the novel and the antique. The conference delivered plenty of results but was also rich in discussion: how should a trial be designed, which endpoints matter most and to whom, and how readily practices lacking true equipoise become normalised in daily care. As a joint piece, Martin and Nel have put together their perspectives from the conference, with one focused on challenging clinical assumptions and choosing appropriate trial designs, and the other on the statistical challenges of generating and interpreting meaningful evidence.

Proving and disproving the interventions implemented in Critical Care - Nel
CCR was fascinating and touched upon plenty of the current challenges faced within critical care research. The one which stuck most with me was that ‘physiological plausibility does not guarantee positive clinical outcomes.’ Despite this being a scientifically straightforward statement, it encapsulates many challenges in its consequences for trial research and clinical practice.
The first consequence of this statement is for procedures which are currently routine in critical care but, rather than backed with evidence, are used due to their physiological plausibility. These live in clinical equipoise but seem not to be contested or questioned by a large part of the community. The GASTRIC-PICU trial looked at one such issue. Designed as a non-inferiority trial, the team found that not measuring Gastric residual Volume (GRV) routinely in children was not worse for survival and mechanical ventilation-free days than at least 6-hourly measurement. Not only that, but some practitioners questioned whether the procedure on adults had any clinical merit too.
This example highlights that many times researchers focus on exciting new treatments or modern ways to understand the aetiology of the disease through biomarkers or genetics. However, further research is still needed to understand some routine current clinical practices, in the context of whether it improves the outcomes of all patients targeted, or only for some patients with a given disease or condition.
On that same line, REMAP-CAP presented stark results, concluding that the use of oseltamivir in critically ill patients with influenza in ICU had less than a 2% probability of being effective and risked being potentially harmful, despite overwhelming evidence that the treatment is beneficial in reducing illness duration in mild/moderate influenza patients. This slight difference in the population is what can turn the treatment from effective to harmful, contradicting the physiologically plausible argument that it could show benefit in all similar populations.
So, what does this mean for how we design trials? Mostly, it makes me reflect on how to get the right research question and design to answer it. Superiority designs suit new treatments; de-implementation research, meaning "can we stop doing this", usually needs a non-inferiority design with an appropriate non-inferiority margin agreed in advance. Conversations around the margin of non-inferiority are based on clinical judgement and are not exclusively statistical. Patients and clinicians should be in the room for these discussions, and as GASTRIC-PICU shows with 4700 participants recruited, can require large sample sizes and careful design around non-inferiority. By contrast, Bayesian platform trials such as REMAP-CAP, whilst presenting complexity in design, use posterior distributions to display the probabilities of efficacy of interventions, or even harm if the treatment effect favours usual care. This design choice lets REMAP-CAP present effectiveness and harm information in a population nobody had separately powered for, and by presenting the probability of effectiveness and probability of harm for a given intervention, it hopefully provides new evidence for clinicians on the use of this intervention for this critically ill patient population.
Whilst clinical trials provide the gold standard for clinical evidence, their findings are not necessarily mechanistic by default. Causal links need to be proven between populations, and even if tested, many trials’ secondary outcomes usually fall under the umbrella of a lack of statistical power. As closely matching and interchangeable as two populations may seem, the results of REMAP-CAP’s and GASTRIC-PICU’s presentation show that quality trials are needed to prove generalisability, especially in clinical conditions where treatment heterogeneity can vary or even reverse the intervention’s effects.
Clinical conference: a statistician’s view - Martin
Like trials in every disease area, critical care trials suffer from recruitment challenges, though the field's characteristics make these challenges unusually acute. It is a broad field that encompasses very frail participants, who will require multiple lines of varied care. These two factors, a heterogeneous population and multifaceted treatment regimens, result in many combinations of possible trials to run, and so smaller populations to find samples from. There are also short recruitment windows (as treatment typically starts immediately) and often many competing trials. The recruitment problem is compounded by the choice of primary outcome measures. Many trials resort to binary outcomes such as mortality, and while they are usually well established and clinically justified, they carry the least information of any clinical scale. Consequently, powering a trial using a binary outcome will result in the largest possible sample size versus alternatives. In response, several speakers highlighted the need for information-rich outcome measures, with ordinal outcomes proposed in the editorial for the MARCH trial. Ordinal outcomes can capture a range of health states and are far more statistically efficient than their binary alternatives. So, not only do they give a better insight into the clinical condition of the patient, but they can also lead to a substantial reduction in the required sample size. We see the field moving in this direction as partial versions, such as days alive and free of organ support, already appear in trials like SODa-BiC and ARISE FLUIDS, though they collapse states that differ clinically. As statisticians, we can work with the recommendations from these talks to make meaningful methodological developments in trial design for these outcomes. This discussion showed that collaboration between clinicians, patients, and statisticians can contribute to more creative solutions to the challenges of recruitment.
The next challenge beyond recruitment is the interpretation of the findings. Unlike drug licensing, where evidential requirements are formally codified, there is no defined threshold for changing practice: the process is diffuse and rests on collective professional judgement. This was evident with the BIHCA trial. The trial failed to show superiority, and the panel was technically right that this did not show the drug does not work: absence of evidence is not evidence of absence. Following on from this, a member of the audience declared that because sodium bicarbonate had not been shown to be inferior, they would continue to use it. Both positions are pedantic: they rest on a logical technicality while setting aside the entirety of the evidence. Not only did sodium bicarbonate show no superiority on the primary outcome; there was no evidence that the pathway of effect is as hypothesised, as there were no significant differences on any of the mechanistic secondary outcomes either. Add to this that BIHCA was a nationwide trial that was exceptionally demanding to run, and you start to ask how far we must go. Ironically, having no fixed evidential threshold only appears to support the Bayesian case against deciding by a single arbitrary cut-off. In practice it did not lead to a careful weighing of the evidence: one clinician simply took a null result as reason to continue using the drug. A Bayesian would say this is because clinicians are still accustomed to interpreting evidence as significant or not, and that this habit will decline over time. This may well be true. But in the intervening period, though uniform criteria are not without problems, their absence leaves practice change resting on interpretations that vary widely between clinicians, as this discussion illustrated. Patients' time is wasted, and vast sums are spent gathering evidence that has little effect on practice. The risk is not only practical but ethical. The discussion over BIHCA added valuable substance to the debate between subjective and objective evaluation of evidence.
The final, and in my view most important, discussion was over patient-reported and clinical outcomes. We were reminded that it is always easier to assume what is right for the patient rather than consulting them properly, and that patient-reported outcomes should have the final say on whether an intervention works or not. Not exactly radical thinking, and I am reassured that, from where I sit, there is something close to consensus on this, but CCR highlighted an instance suggesting otherwise. Referring back to the BIHCA trial, whose primary outcome was sustained return of spontaneous circulation (ROSC), the evidence showed no difference on this outcome, and this was subsequently taken as the final word on the superiority of sodium bicarbonate, since without ROSC death is certain. One can sympathise with the investigators: how can this be controversial, given that ROSC is a prerequisite for survival? I myself am quite fond of being alive. The issue is that ROSC is necessary but not sufficient for an evaluation of the treatment. Yes, not being dead is of utmost interest to the patient, but for those who do survive, we must ask what their quality of life is and what the health implications are. These are vital questions, and conducting a trial on surrogate endpoints alone leaves us ignorant of them. This issue was raised by the panel, who recommended that this additional research be undertaken. If patient-reported outcomes are to have the final say, our endpoints must first put the question to them.
As a statistician, I am in the habit of thinking in one direction when translating clinical questions into statistical ones. CCR argued the process is more iterative. Statistics must serve the messy realities of medical research; the textbook case never arrives. Compromises must be made, but out of these, new insights and new developments can emerge.
Closing remarks
Attending CCR exposed us to the many challenges currently faced in the critical care setting. We discussed the importance of clinical evidence over physiological theory, the choice of outcome measures, patient-centred outcomes, and the interpretation of uncertain findings. These themes stem from a general principle worth repeating: clinical trial design must revolve around patients, clinicians and clinical practice. Whether it is the choice of design, the primary outcome measure or the interpretation of uncertain findings, this holds not just for critical care but for clinical trials in general.
Key takeaways:
Physiological plausibility does not guarantee positive clinical outcomes
Design, outcomes, and inferential framework must revolve around patients, clinicians and clinical practice
Statistical innovation does not happen in isolation and is an iterative process with clinicians and patients
References and resources
All the information on the trials and the editorials at CCR can be found here: https://criticalcarereviews.com/meetings/ccr26



Comments