Regulatory ImpactRegulatory Impact
Regulatory InsightsClinical Development

Novel Clinician-Reported Outcomes in FDA Approvals: Five Regulatory Lessons Since 2020

Regulatory ImpactJuly 30, 202614 min read
Clinical Outcome AssessmentsClinician-Reported OutcomesFDAEndpoint StrategyRare DiseaseRegulatory Review

Executive Summary

The U.S. Food and Drug Administration has approved products whose efficacy packages relied on clinician-reported outcomes in markedly different ways. Some instruments served as primary endpoints, some contributed one component of a composite responder definition, and others supplied supportive evidence that connected an objective biological response to a functional benefit.

The approvals do not establish a single pathway for novel clinician-reported outcomes. They show a recurring regulatory pattern instead: the agency evaluates the measured concept, the observer, the scoring system, the administration procedures, the statistical estimand, and the proposed labeling interpretation as one integrated endpoint system.

The central lesson is that instrument acceptance is not equivalent to evidence sufficiency. A scale may be considered capable of measuring a clinically relevant concept while the application still faces questions about effect size, rater reliability, domain weighting, missing data, meaningful change, multiplicity, or the need for confirmatory evidence. Conversely, an imperfect instrument may remain usable when a narrowly defined endpoint is inherently interpretable, the result is robust, and the totality of evidence supports the intended claim.

What Makes a Clinician-Reported Outcome Novel

The U.S. Food and Drug Administration (FDA) describes a clinical outcome assessment (COA) as a measure of a patient's symptoms, functioning, or survival. A clinician-reported outcome (ClinRO) is based on a trained healthcare professional's observation and interpretation of observable signs, behaviors, or other clinical manifestations. It differs from a patient-reported outcome (PRO), which comes directly from the patient without interpretation by another person, and from a performance outcome (PerfO), which is based on a standardized task performed by the patient. 1

For this analysis, novel is an operational description rather than an FDA classification. It includes newly developed scales, materially modified instruments, newly selected domains or subscores, and established clinical measurements placed into a new endpoint construction or context of use. FDA's patient-focused drug development guidance frames COA selection, development, and modification as a fit-for-purpose exercise, meaning that acceptability depends on the concept, population, trial design, administration method, analysis, and intended interpretation. 2

FDA's COA Compendium can help sponsors identify measures previously used in drug development, but inclusion is not an endorsement and does not substitute for engagement with the relevant review division. The post-2020 cases examined here show why: review questions repeatedly concerned content validity, rater training, reliability, scoring rules, anchors for meaningful change, missing data, endpoint hierarchy, and the relationship between the endpoint result and the proposed label. 34567

Selected Approvals at a Glance

The following cases are a selected, non-exhaustive set of approvals from 2020 through 2025. They were chosen because the public FDA reviews provide unusually clear examples of how a novel or modified ClinRO affected endpoint acceptability, statistical interpretation, substantial-evidence assessment, or labeling. 89101112

ProductApproval yearClinRO roleCentral review issueRegulatory outcome
Qwo2020Clinician cellulite scale combined with a patient scale in the primary responder definitionAgreement, reliability, standardized image and live assessment, and interpretability of a stringent dual-response thresholdTwo pivotal trials met the multicomponent primary endpoint, although FDA characterized the absolute response rates and treatment effects as small
Spevigo2022Physician global assessment pustulation subscore as the primary endpointWeak evidence for the total score and limited measurement-property workComplete pustular clearance at Week 1 was considered inherently meaningful and supported approval
Skyclarys2023Modified neurological rating scale as the primary efficacy endpointWhether one study with modest effect size and negative key secondary endpoints constituted substantial evidenceFDA accepted the scale as a direct measure and relied on the primary result plus confirmatory evidence in the totality
Miplyffa2024Rescored four-domain disease severity scale as the primary endpointDeficient cognition and swallowing domains, scoring, standardization, estimand, and confirmatory evidence after an initial complete responseThe revised instrument and additional evidence supported approval only in combination with miglustat
Romvimza2025Goniometer-measured active range of motion as a multiplicity-controlled supportive endpointLimited reliability evidence and no well-supported responder thresholdContinuous and individual-patient functional results were included in labeling as support for clinical benefit

Qwo: Validation Before Pivotal Execution

Qwo's pivotal program for cellulite used two parallel five-point photonumeric scales: the Clinician-Reported Photonumeric Cellulite Severity Scale (CR-PCSS) and the Patient-Reported Photonumeric Cellulite Severity Scale (PR-PCSS). The primary endpoint required the same treated buttock to improve by at least two severity levels on both scales at Day 71, creating a stringent multicomponent responder definition that demanded agreement between clinician and patient perspectives. 48

The endpoint was not accepted merely because both concepts appeared clinically relevant. During a 2015 Type C interaction, FDA raised concerns about differences between clinician and patient assessment and about data-collection methods. The sponsor subsequently generated qualitative and quantitative evidence, used standardized photographs and training, and evaluated intra-rater and inter-rater reliability. FDA's review concluded that content validity was reasonably established and that the reliability evidence was sufficient to support the proposed two-level responder threshold, after earlier Special Protocol Assessment concerns had required additional analyses. 4

The two phase 3 trials met the primary endpoint. In one trial, the two-level composite response was 7.6 percent with Qwo and 1.9 percent with placebo, an adjusted difference of 5.7 percentage points with a p-value of 0.006. In the other, response was 5.6 percent and 0.5 percent, respectively, an adjusted difference of 5.1 percentage points with a p-value of 0.002. FDA considered the absolute response rates and treatment effects small, but found the results robust to tipping-point analyses and supported by the separate clinician, patient, and secondary endpoint results. 8

The regulatory value of the Qwo example lies in endpoint architecture. Pairing two perspectives can strengthen the clinical meaning of a response, but it also compounds measurement error and lowers response rates because both components must be satisfied. Sponsors using this structure need evidence for each instrument, a prespecified rule for combining them, training that minimizes avoidable discordance, and sensitivity analyses that show the conclusion is not an artifact of missing data or a single component. 48

Spevigo: The Subscore Survived the Total Score

Spevigo was evaluated in adults with an acute generalized pustular psoriasis (GPP) flare using the Generalized Pustular Psoriasis Physician Global Assessment (GPPPGA). The instrument contained separate pustulation, erythema, and scaling components, but the pivotal primary endpoint was the proportion of patients with a pustulation subscore of zero at Week 1. In the randomized trial, 54.3 percent of patients assigned to Spevigo and 5.6 percent assigned to placebo achieved that endpoint, with a p-value of 0.0004. 59

FDA's COA review distinguished sharply between the total GPPPGA score and the pustulation subscore. The agency found the total-score interpretation inadequate because the observed improvement was largely driven by pustulation. It also identified gaps in the development program, including the absence of qualitative clinician interviews, no inter-rater reliability assessment, a small sample, and anchors that did not adequately support interpretation of score change. 5

Those deficiencies did not invalidate the primary endpoint. A score of zero on the pustulation domain represented complete absence of visible pustules, so FDA considered the endpoint inherently meaningful even though anchor-based analyses were not interpretable. The agency concluded that the pustulation subscore could support labeling, while recommending more systematic measurement-property work for future development. 5

Spevigo illustrates the importance of separating an instrument from a particular scoring use. A multidomain total score can obscure which manifestation is changing, especially when domains differ in clinical salience or responsiveness. A prespecified, clinically coherent subscore may be more defensible than a total score when it measures the defining acute manifestation and its endpoint state has a direct interpretation such as complete clearance. 59

Skyclarys: Endpoint Acceptance Was Not Evidence Sufficiency

Skyclarys was developed for Friedreich's ataxia (FA) using the modified Friedreich's Ataxia Rating Scale (mFARS) as the pivotal primary endpoint. The modification removed section D of the original scale, which FDA and the sponsor considered less clinically meaningful, while retaining neurological examination domains in which higher scores indicate greater impairment. Regulatory discussions evolved over several years: FDA initially recommended an activities of daily living measure as primary, later accepted mFARS, requested supportive patient-reported or performance-based measures, and supported extending the controlled study to 48 weeks. 13

In the prespecified full analysis set of 82 patients without pes cavus, the placebo-adjusted difference in mFARS change at Week 48 was minus 2.41 points, with a p-value of 0.0138. The omaveloxolone group improved by a mean of 1.56 points while placebo worsened by 0.85 points. The key secondary endpoints were not statistically significant in that analysis population, and FDA had previously advised that a single study with the observed effect size, statistical strength, and secondary-endpoint pattern would not by itself be sufficient. 1310

At approval, FDA treated mFARS as an acceptable direct measure of neurological function but evaluated substantial evidence separately. The agency considered the pivotal result together with additional evidence, including pharmacodynamic information from an earlier study and a comparison with natural-history data. FDA concluded that the totality supported effectiveness for patients 16 years of age and older, while acknowledging limitations in the supporting analyses. 10

The Skyclarys review shows why a favorable endpoint discussion should not be mistaken for agreement on the evidentiary package. Sponsors need two linked strategies: one that establishes that the ClinRO measures an important concept reliably and interpretably, and another that establishes how many trials, what statistical strength, and what corroborating evidence will be required if the pivotal result is modest or secondary endpoints are not persuasive. 1310

Miplyffa: Repairing a ClinRO After Complete Response

Miplyffa provides the clearest example in this set of a ClinRO being materially reworked after an unsuccessful first review cycle. The original application used a five-domain Niemann-Pick disease type C Clinical Severity Scale (5DNPCCSS) covering ambulation, cognition, fine motor skills, speech, and swallowing. FDA issued a complete response letter in June 2021 that identified deficiencies in endpoint interpretability, the prespecified analysis, and the proposed confirmatory evidence. The agency's COA concerns included a cognition domain that was not fit for purpose, overlapping or poorly ordered swallowing response options, possible silent aspiration, and insufficient standardization across several domains. 611

At an October 2021 Type A meeting, FDA and the applicant agreed to remove cognition. The applicant also revised swallowing scoring using additional clinical input and a qualitative clinician study, producing the rescored four-domain Niemann-Pick disease type C Clinical Severity Scale (R4DNPCCSS), which retained swallowing, speech, fine motor skills, and ambulation. FDA continued to identify limitations in standardization and reliability, but considered the revised measure interpretable for this application in light of the disease context and unmet need. 6

The original five-domain analysis produced an estimated treatment effect of minus 1.4 points with a p-value of 0.0456. The revised four-domain while-on-treatment analysis also estimated a minus 1.4-point effect, with a p-value of 0.0451, while the treatment-policy confidence interval crossed zero. Approximately 78 percent of randomized patients received background miglustat. In the miglustat subgroup, the estimated treatment difference was minus 2.2 points with a p-value of 0.0028, but the smaller subgroup without miglustat could not establish an effect of arimoclomol alone. 11

The approved labeling therefore narrowed the evidentiary interpretation to the setting supported by the data. Miplyffa was approved in combination with miglustat for neurological manifestations of Niemann-Pick disease type C in adults and children two years of age and older. The label reports a placebo-adjusted difference of minus 2.2 points on the R4DNPCCSS in patients receiving miglustat and states that there were insufficient data to determine effectiveness without miglustat. 14

The Miplyffa sequence demonstrates that post hoc endpoint repair is possible, but it is not a simple rescoring exercise. The sponsor had to connect each change to concept relevance, clinician interpretation, scoring behavior, estimand choice, sensitivity analyses, and the broader evidence package. The eventual indication also reflected the subgroup in which the effect was interpretable, showing how unresolved endpoint and analysis uncertainty can propagate directly into the approved use. 61114

Romvimza: Supportive Function Without a Settled Threshold

Romvimza was studied in adults with symptomatic tenosynovial giant cell tumor (TGCT). The pivotal trial's primary endpoint was objective response rate (ORR) by Response Evaluation Criteria in Solid Tumors version 1.1, while active range of motion (ROM), measured by a clinician using a goniometer, was included as a multiplicity-controlled secondary endpoint. At Week 25, the least-squares mean difference in active ROM was 14.6 percentage points of the normal reference range, with a 95 percent confidence interval from 4.0 to 25.3 and a p-value of 0.0077. 1215

FDA's COA review considered active ROM relevant to the functional consequences of TGCT but found the measurement-property package incomplete. The review identified insufficient evidence for goniometer reliability in the trial context and insufficient support for a meaningful within-patient change threshold. The proposed anchors had conceptual and statistical limitations, and the patient global impression anchor for range of motion had 24.4 percent missing data at Week 25. These issues had been preceded by FDA requests across Type C and end-of-phase 2 (EOP2) interactions for evidence on goniometry, scoring, anchors, missing data, and cumulative distribution displays. 7

The absence of a validated responder threshold did not eliminate the endpoint's regulatory contribution. FDA treated active ROM as supportive evidence alongside the tumor response result and other clinical outcome assessments. The prescribing information presented the group-level result and individual-patient changes in active ROM rather than relying on a dichotomized responder claim. 1215

Romvimza demonstrates a viable fallback pathway for a clinically relevant ClinRO when meaningful-change evidence remains unsettled. A continuous endpoint can still support benefit if the analysis is prespecified and multiplicity controlled, the direction and magnitude are interpretable, individual results are transparent, and the functional finding converges with the objective primary endpoint. This is a more limited claim than a validated responder statement, but it can materially improve the clinical interpretation of an otherwise anatomical efficacy result. 71215

Regulatory Design Implications

Across these approvals, the regulatory unit is not the scale alone. It is the full endpoint system: concept of interest, context of use, rater qualification, administration instructions, scoring algorithm, assessment schedule, estimand, handling of intercurrent events and missing data, statistical hierarchy, interpretation threshold, and proposed labeling language. Weakness in any one component can narrow the endpoint's role even when the underlying clinical concept is important. 24567

Design decisionFailure mode seen in reviewRegulatory control
Select the measured concept and domainsA total score is driven by one responsive domain, or includes a domain that is not fit for purposeTest domain contribution before pivotal use; prespecify clinically coherent domain endpoints or exclusions
Build the rater systemTraining exists, but reliability or standardization is not demonstrated in the intended settingUse detailed manuals, certification, calibration, retraining, drift monitoring, and trial-context reliability studies
Define meaningful changeAnchors are poorly aligned, weakly correlated, missing, or difficult to interpretEstablish conceptual alignment through qualitative work; collect multiple fit-for-purpose anchors; preserve a continuous-analysis pathway
Specify the endpoint and estimandPost hoc rescoring or alternative handling of intercurrent events becomes central to the resultFinalize scoring, estimand, intercurrent-event rules, missing-data methods, and sensitivity analyses before unblinding
Design the evidence hierarchyThe primary ClinRO succeeds, but effect size or secondary evidence is not persuasivePrespecify corroborating patient-reported, performance, biological, or external evidence appropriate to the disease and claim
Plan the labeling fallbackA responder threshold is not supported even though the measure is clinically relevantConsider multiplicity-controlled continuous results, distributional displays, individual-patient change, and convergence with objective endpoints

FDA interactions should progressively convert that system into written agreements or clearly documented areas of residual risk. Early meetings can address whether the concept and observer are appropriate. EOP2 and protocol-focused interactions should resolve instrument version, domain selection, rater controls, endpoint definition, estimand, multiplicity, missing-data strategy, supportive measures, and the intended label interpretation. The case histories show that concerns raised before pivotal execution can remain decisive at review when the resulting evidence does not directly answer them. 41367

A sponsor should also maintain a traceable COA evidence dossier rather than distributing the rationale across protocols, manuals, psychometric reports, meeting packages, and statistical plans. The dossier should connect patient and clinician input to item content, define scoring and administration, document training and reliability, justify anchors and thresholds, specify estimands and sensitivity analyses, and map each endpoint use to the exact claim it is intended to support. 2567

Conclusion

Novel clinician-reported outcomes have supported FDA approvals through more than one evidentiary route. They have operated as components of primary responder definitions, focused disease-domain endpoints, modified neurological scales, revised rare-disease severity measures, and supportive functional assessments. Their regulatory value has depended less on novelty itself than on the precision with which the endpoint was defined and connected to a clinically meaningful treatment effect.

The strongest development strategy treats the outcome measure, analysis, and intended claim as a single design problem. Sponsors should establish the relevance of the concept, control observer variability, prespecify how the score will be used, and build an evidence hierarchy that remains persuasive if the primary result is modest or the preferred responder interpretation is not supported. That approach does not eliminate regulatory uncertainty, but it prevents avoidable measurement uncertainty from becoming the central question at application review.