Sponsors who want to use digitally derived endpoints from digital health technologies (DHT) should validate those endpoints for the intended patient population, pre-specify their missing-data strategy at every level of their clinical trial, and avoid conflating device clearance with endpoint validation, according to US Food and Drug Administration (FDA) officials who spoke at a recent public meeting.
The Duke-Margolis Institute for Health Policy and FDA held a meeting on 27 August to discuss statistical considerations for digitally derived endpoints in clinical trials for drug and biological products. The meeting was held a week after FDA published a white paper on key considerations for sponsors using digitally derived measures (DDM) based on existing guidances. The agency noted that DDMs may be used as clinical outcome assessments (COAs), biomarkers, or components of multicomponent endpoints derived from multimodal data, and that their use in trials should be both clinically relevant and meaningful to patients.
During the meeting, Cynthia Fisher, a statistician in the Division of Analytics and Informatics at the Center for Drug Evaluation and Research, discussed the agency’s experience reviewing submissions with digitally derived endpoints in clinical trials. She noted that the information was based on product reviews using a variety of regulatory pathways, including traditional investigational new drug, new drug approval, and biologics license applications, as well as through the Drug Development Tools (DDT) Qualification Program and the innovative Science and Technology Approaches for New Drugs (iSTAND) Program.
Fisher noted that the statistical and methodological issues they have seen with the use of digitally derived endpoints proposed by sponsors are similar regardless of the pathway stakeholders use and offered recommendations on how to avoid such issues.
Fisher said one of the most common deficiencies reviewers have seen is the lack of algorithm validation in the target population. She emphasized that algorithms must be validated in the intended study populations, including variables such as disease severity, age, and movement characteristics of those being monitored.
"All of these things affect sensor signal and algorithm performance in ways that can be substantial and clinically meaningful, but what we frequently receive instead are algorithms that are validated in healthy volunteers or in populations that are demographically distinct from the intended study population," said Fisher. "And we sometimes see key subgroups, like specific age ranges or disease severity strata, actually being excluded from the training data entirely.
"The assumption appears to be that if the algorithm works well in one group, then it generalizes,” she added. “But that assumption is not supportable without evidence."
Fisher said that such misclassification during the algorithm validation process may bias the endpoint in ways that are extremely difficult to detect or correct after the fact, so that by the time regulators realize there's a problem with the pivotal trial, the damage is already done.
"It's important to understand that validation isn't this one-time event that sort of travels with the device; it's population-specific, and it needs to be demonstrated in the people you're actually studying," she added.
Fisher presented some of the most common feedback from reviewers when they find such issues with the proposed digitally-derived endpoints presented by sponsors, including that their training dataset does not adequately represent the intended clinical population, that an independent algorithm validation study in the target population is recommended before sponsors begin their pivotal trial, and that it is unclear to reviewers whether the sponsor's proposed algorithm is valid in the target population.
Another common issue reviewers face is that sponsors often fail to fully pre-specify a missing data strategy. Fisher said they must pre-specify a missing-data strategy at every level of the data hierarchy and noted that submissions commonly address only the subject-level strategy while leaving strategies at lower levels unspecified.
Lower-level missing data strategies are important, according to Fisher, because regulators need to understand when and why data may be missing. Besides being told by reviewers that they should specify missing-data handling at each level of data aggregation, she said reviewers may also tell sponsors that the assumption that data are missing at random may not be viable in their trial population and that they should provide sensitivity analyses under alternative missing-data mechanisms. Furthermore, they may be asked to evaluate the impact of different missingness thresholds on the endpoint estimates.
Fisher said that a small amount of missing data that is truly random, such as that caused by a device malfunction or due to a transmission error, is generally manageable. However, she noted that missing data is rarely random, which is why it's important to understand its reasons.
"The question, I'd encourage sponsors to ask it not, 'How much data is missing,' but 'Who is missing it, and when are they missing it, and why are they missing it,'" said Fisher. "That means actively collecting that information.
"That could be device error codes, subject diaries, site logs, and build that kind of stuff into your protocol, because that is what will help tell you if there is enough information to trust your results," she added. "What makes us more comfortable, even with moderate amounts of missing data, is prespecified sensitivity analyses."
Fisher said that another challenge reviewers face is patient-centered justification for endpoints. She said that endpoint selection, including the specific metric and threshold, must be based on patient input rather than on what is analytically convenient or on the DHT device's capabilities.
"This is where usability testing or cognitive interviews are really important because they help confirm that patients can actually engage with the device as intended in their daily lives and that what the device captures reflects their experience," said Fisher.
If sponsors do not sufficiently provide patient-centered justifications for their endpoints, Fisher said reviewers may ask sponsors to go back to provide qualitative evidence showing their endpoints are patient-centered, that multiple metrics are derivable from their monitoring devices while providing the rationale behind choosing specific metrics, and provide usability testing with cognitive interviews to confirm patients can use the devices as intended.
Fisher said sponsors often conflate device clearance with endpoint validity.
"First, I think it's important to state that a device does not need to be 510k cleared to be used in a clinical trial," said Fisher. "Secondly, even if you are using a cleared device, clearance and endpoint validity are separate regulatory questions.
"Clearance establishes that a device measures something accurately under the conditions and indications of that cleared submission, nothing more," she added. "It doesn't establish that the device is fit-for-purpose as the basis of an efficacy endpoint in a different population, under a different protocol, using a different clinical construct."
Fisher noted that the device and the endpoint it generates must each independently demonstrate that they are fit for purpose, and that demonstration must be made in the context of the product development program. She also noted that unless the populations being investigated are sufficiently similar, sponsors can't automatically transfer endpoint validation from a prior clinical trial to a new one, because endpoints must independently meet the evidentiary threshold, depending on the circumstances of the trials.
"Whether the device can measure something is only part of the question; the other part is whether what it measures can support a regulatory conclusion," said Fisher. "The good news is every gap I've described here is preventable.
"The sponsors who avoid these deficiencies aren't doing anything extraordinary," she added. "They're validating in the right population, they're prespecifying at every level, and they're anchoring their endpoints in patient experience, and perhaps most importantly... they're engaging with FDA early."
Fisher emphasized that early engagement with the agency allows sponsors more time and flexibility to address potential deficiencies.
“Engage with us earlier than you think you need to,” said Fisher. “The sponsors who avoid the issues that we've all been mentioning are the ones who are coming to the table early, coming to the table when there's still time to course correct.
“We see that when a full submission arrives that has a fundamental gap in algorithm validation or endpoint justification, the options for fixing it are very limited and expensive,” she added. “When you have early engagement, we can help raise those concerns when there's still time and flexibility.”
Jessica Zhou, a digital health specialist at FDA, added that it's important to be prepared more broadly. In particular, she said sponsors should prepare their DHT data for regulatory use early in the product development. She also said that, in addition to ensuring the device and data are fit for purpose, it's important to determine how much data is needed and set thresholds for variables such as wear time and signal quality.
"It's those thresholds that can support alerts that can identify problems early and really allow sites to troubleshoot or provide device replacements," said Zhao. "What we want to avoid is discovering late in the trial that a device malfunctioned or the validation was insufficient or there's too little usable data collected."
Fisher noted several challenges reviewers encounter when evaluating submissions with digitally derived endpoints, such as missing data because patients didn't charge their DHT devices, removed their devices during their worst symptom days, or wore their devices incorrectly, which "can be hard to recover from analytically." Another challenge she highlighted is inconsistency across clinical trial sites regarding issues such as standardized training on device placement, charging schedules, and troubleshooting.
"All of these aspects are very important but often under documented," said Fisher. "We see that when sites handle things differently, you introduce variability that may look like a signal… but it's really just noise."
"If you can build those solutions into your protocol before you enroll any patients, you can avoid a lot of these problems down the road," she added.