<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "journalpublishing.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="2.0" xml:lang="en" article-type="article-commentary"><front><journal-meta><journal-id journal-id-type="nlm-ta">J Med Internet Res</journal-id><journal-id journal-id-type="publisher-id">jmir</journal-id><journal-id journal-id-type="index">1</journal-id><journal-title>Journal of Medical Internet Research</journal-title><abbrev-journal-title>J Med Internet Res</abbrev-journal-title><issn pub-type="epub">1438-8871</issn><publisher><publisher-name>JMIR Publications</publisher-name><publisher-loc>Toronto, Canada</publisher-loc></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">v28i1e111711</article-id><article-id pub-id-type="doi">10.2196/111711</article-id><article-categories><subj-group subj-group-type="heading"><subject>Commentary</subject></subj-group></article-categories><title-group><article-title>Explicit Mechanistic Causal Analyses or Interventional Trials Are Required for Objective, Clinical, Voice-Based Parkinson Disease Characterization</article-title></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name name-style="western"><surname>Little</surname><given-names>Max</given-names></name><degrees>DPhil</degrees><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff id="aff1"><institution>School of Computer Science, University of Birmingham</institution><addr-line>Edgbaston</addr-line><addr-line>Birmingham</addr-line><addr-line>England</addr-line><country>United Kingdom</country></aff><contrib-group><contrib contrib-type="editor"><name name-style="western"><surname>Schwartz</surname><given-names>Amy</given-names></name></contrib><contrib contrib-type="editor"><name name-style="western"><surname>Leung</surname><given-names>Tiffany</given-names></name></contrib></contrib-group><author-notes><corresp>Correspondence to Max Little, DPhil, School of Computer Science, University of Birmingham, Edgbaston, Birmingham, England, B15 2TT, United Kingdom, 44 121 414 3344; <email>maxl@mit.edu</email></corresp></author-notes><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>28</day><month>9</month><year>2026</year></pub-date><volume>28</volume><elocation-id>e111711</elocation-id><history><date date-type="received"><day>10</day><month>09</month><year>2026</year></date><date date-type="rev-recd"><day>14</day><month>09</month><year>2026</year></date><date date-type="accepted"><day>16</day><month>09</month><year>2026</year></date></history><copyright-statement>&#x00A9; Max Little. Originally published in the Journal of Medical Internet Research (<ext-link ext-link-type="uri" xlink:href="https://www.jmir.org">https://www.jmir.org</ext-link>), 28.9.2026. </copyright-statement><copyright-year>2026</copyright-year><license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (<ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in the Journal of Medical Internet Research (ISSN 1438-8871), is properly cited. The complete bibliographic information, a link to the original publication on <ext-link ext-link-type="uri" xlink:href="https://www.jmir.org/">https://www.jmir.org/</ext-link>, as well as this copyright and license information must be included.</p></license><self-uri xlink:type="simple" xlink:href="https://www.jmir.org/2026/1/e111711"/><related-article related-article-type="commentary article" ext-link-type="doi" xlink:href="10.2196/95609" xlink:title="Comment on" xlink:type="simple">https://www.jmir.org/2026/1/e95609</related-article><abstract><p>For most with Parkinson disease, a movement disorder, voice and speech are impaired at some point, which presents an opportunity to use digital recordings and sophisticated machine learning to assist objective clinical characterization of the condition. However, as highlighted by Shukla et al in their August 20, 2026 paper, ad-hoc observational datasets typically contain spurious causal associations that escape a purely statistical analysis. This commentary argues that causal inference, and ultimately, diagnostic clinical trials, are required to establish a direct mechanistic relationship between disease and algorithm predictions.</p></abstract><kwd-group><kwd>digital voice analysis</kwd><kwd>causal inference</kwd><kwd>diagnostic trials</kwd><kwd>machine learning</kwd><kwd>confounding</kwd></kwd-group></article-meta></front><body><p>Shukla et al [<xref ref-type="bibr" rid="ref1">1</xref>] demonstrate that naive statistical analysis of digital recordings to assist objective, clinical, voice-based characterization of Parkinson disease can be highly misleading, even with sophisticated machine learning algorithms. However, this is a consequence of ad hoc observational datasets, which typically contain spurious causal associations that escape a purely statistical analysis. Causal inference, and ultimately, diagnostic clinical trials, are required to establish a direct mechanistic relationship between disease and algorithm predictions.</p><p>There has been considerable interest, stretching back 30 or more years, in using voice and/or speech for objective clinical characterization of certain disorders and diseases that cause vocal signs and symptoms [<xref ref-type="bibr" rid="ref2">2</xref>]. Sounds are readily captured noninvasively by ubiquitous digital recording equipment, and recent progress in machine learning (ML) algorithms has enabled nearly fully automated processing of these digital signals to predict clinical labels such as positive/negative diagnoses or symptom severity by training these algorithms on cohort datasets (<xref ref-type="fig" rid="figure1">Figure 1A</xref>). Movement disorders such as Parkinson disease are prime examples of clinical conditions that are inherently amenable to this approach.</p><fig position="float" id="figure1"><label>Figure 1.</label><caption><p>(A) Naive machine learning prediction of clinical disease status (X) from digital voice/speech features (Y) assumes features are solely dependent upon the underlying disease process and nothing else. (B) The situation is generally more structurally complex, however: age, recording environment, and/or identity are very often common causes of both the digital recordings and disease status, so they are causal confounders (Z). (C) Causal inference can break the spurious association path this entails. (D) Nevertheless, causal inference is not always possible, for instance, where there is hidden confounding (H). (E) In that situation, only an interventional trial can obtain a scientifically defensible estimate of machine learning prediction performance.</p></caption><graphic alt-version="no" mimetype="image" position="float" xlink:type="simple" xlink:href="jmir_v28i1e111711_fig01.png"/></fig><p>Nonetheless, a naive application of these algorithms overlooks common statistical issues with these observational datasets, which are typically collected ad hoc absent principled experimental design protocols. For instance, there may be imbalanced ages or differences in acoustic environments between cases and controls. These imbalances can be sufficiently large that a simple benchmark prediction based on age or environment can be more accurate than predictions based on the digital audio itself. Shukla et al [<xref ref-type="bibr" rid="ref1">1</xref>] demonstrate exactly this phenomenon with a specific dataset (Bridge2AI-Voice), using a sophisticated deep learning transformer architecture applied to the digital audio features.</p><p>This is clearly scientifically dubious because objective characterization should be based on the audio recordings of the patients, not on irrelevances such as patient age or acoustic setting. Bayesian analysis can be used for rebalancing via a suitable choice of prior, but it is specific to the particular dataset [<xref ref-type="bibr" rid="ref3">3</xref>]. Majority/minority resampling methods can compensate for the imbalance by modifying the data itself. These either repeat samples, introducing statistical artifacts, or discard data, which reduces the available data size and thereby increases the statistical uncertainty of estimation for these ML models.</p><p>However, the problems above, arising from ad hoc data collection settings, often run deeper. Age can affect voice and speech (presbyphonia), and many diseases and syndromes become more common with advancing age. Age can thus be a common cause of both clinical label and vocal features [<xref ref-type="bibr" rid="ref4">4</xref>]. Similarly, individuals have distinct vocal identities, and many diseases of interest are chronic or progressive and do not undergo spontaneous remission, so datasets may have no recordings of individuals in both control and case state, even when they contain multiple recordings from the same individual [<xref ref-type="bibr" rid="ref5">5</xref>]. Thus, particularly in small cohorts, both clinical label and vocal features have another common cause: the individual&#x2019;s uniqueness. An algorithm must detect the <italic>direct</italic> causal relationship between vocal signs and symptoms and clinical label, but these nuisance causes produce <italic>spurious</italic> associations that easily swamp sensitive ML algorithms [<xref ref-type="bibr" rid="ref5">5</xref>].</p><p>This mechanistic framing gives a rigorous structural causal meaning (<xref ref-type="fig" rid="figure1">Figure 1B</xref>) to confounding, loosely used by Shukla et al [<xref ref-type="bibr" rid="ref1">1</xref>]. Causal problems such as this are, in general, provably beyond the reach of purely statistical analysis [<xref ref-type="bibr" rid="ref4">4</xref>]; indeed, cross-validation (CV) [<xref ref-type="bibr" rid="ref6">6</xref>] cannot correct for causal confounding [<xref ref-type="bibr" rid="ref5">5</xref>]. Stratified (individual-level) CV [<xref ref-type="bibr" rid="ref1">1</xref>,<xref ref-type="bibr" rid="ref6">6</xref>] attempts to resolve issues such as identity, age, or sex confounding, but this generally biases the quantification of prediction error because, by construction, the training and test sets no longer have the same distribution, which invalidates hold-out methods like CV [<xref ref-type="bibr" rid="ref3">3</xref>,<xref ref-type="bibr" rid="ref5">5</xref>,<xref ref-type="bibr" rid="ref6">6</xref>].</p><p>Solving causal confounding in these datasets can be achieved through explicit causal inference methods (<xref ref-type="fig" rid="figure1">Figure 1C</xref>), the simplest of which (applicable in the backdoor case) is controlling for confounders by including them as covariates in the prediction model [<xref ref-type="bibr" rid="ref4">4</xref>]. Covariate controlling is unsatisfactory, however, because the algorithm then requires the value of the confounder at prediction time. Probabilistic adjustment methods remove this restriction, but this comes at the expense of requiring an explicitly probabilistic model of all joint variables in the problem [<xref ref-type="bibr" rid="ref4">4</xref>]. Many sophisticated ML algorithms (such as deep learning) are not explicitly probabilistic, and modeling high-dimensional digital audio data is difficult. Shukla et al [<xref ref-type="bibr" rid="ref1">1</xref>] use propensity score matching, applicable in the backdoor case, which constructs a subsample that approximately mimics the causal consequences of removing the association between confounders and clinical label (<xref ref-type="fig" rid="figure1">Figure 1C</xref>). However, discarding data like this always increases statistical uncertainty and demands positivity, which is violated in deterministic confounding such as the spurious identity/vocal uniqueness association and in small datasets in general [<xref ref-type="bibr" rid="ref7">7</xref>,<xref ref-type="bibr" rid="ref8">8</xref>]. If such adjustment is possible then counterfactual inference may be used to test more complex predictions beyond the scope of purely interventional causal logic, such as hypothetical case/control assignments for specific individuals [<xref ref-type="bibr" rid="ref4">4</xref>].</p><p>Increasingly sophisticated causal machine learning is one route to more scientifically defensible ML modeling in this setting, but causal inference demands sufficiently precise knowledge of the structural causal mechanisms involved. There can be hidden factors (<xref ref-type="fig" rid="figure1">Figure 1D</xref>) such that only in certain causal arrangements is the direct effect identifiable from the observational data [<xref ref-type="bibr" rid="ref4">4</xref>]. Thus, ultimately, explicitly interventional data collection (eg, through randomized controlled trials) is required to completely eliminate bias due to spurious associations (<xref ref-type="fig" rid="figure1">Figure 1E</xref>). However, the usual problem to solve in this setting is predicting clinical labels or scores, and it is impossible (or unethical) to impose this status on trial participants as an intervention. Thus, we must settle for diagnostic trials [<xref ref-type="bibr" rid="ref9">9</xref>] where treatment outcome is generally used as the effect, and the intervention is the use of the novel voice/speech-based prediction algorithm versus the current standard-of-care.</p><p>Clinical trials like this are, in theory, the gold standard test of digital voice/speech-based symptom prediction using ML, but only in terms of treatment outcomes. This is fundamentally not achievable with typical ad hoc observational datasets used in ML studies of the kind that the discipline has so far examined. Diagnostic trials are thus, very likely, the end goal for this line of digital diagnostic assistance technology, if it is to be scientifically reliable for real-world clinical deployment.</p></body><back><ack><p>No use was made of generative AI in preparing this commentary.</p></ack><notes><sec><title>Funding</title><p>The author declares no financial support was received for this work.</p></sec></notes><fn-group><fn fn-type="conflict"><p>None declared.</p></fn></fn-group><glossary><title>Abbreviations</title><def-list><def-item><term id="abb1">CV</term><def><p>cross-validation</p></def></def-item><def-item><term id="abb2">ML</term><def><p>machine learning</p></def></def-item></def-list></glossary><ref-list><title>References</title><ref id="ref1"><label>1</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Shukla</surname><given-names>S</given-names> </name><name name-style="western"><surname>Naliyatthaliyazchayil</surname><given-names>P</given-names> </name><name name-style="western"><surname>Gichoya</surname><given-names>JW</given-names> </name><name name-style="western"><surname>Purkayastha</surname><given-names>S</given-names> </name></person-group><article-title>Demographic confounding in voice-based Parkinson disease screening: methodological analysis of the Bridge2AI voice dataset</article-title><source>J Med Internet Res</source><year>2026</year><month>08</month><day>20</day><volume>28</volume><fpage>e95609</fpage><pub-id pub-id-type="doi">10.2196/95609</pub-id><pub-id pub-id-type="medline">42623305</pub-id></nlm-citation></ref><ref id="ref2"><label>2</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Boyanov</surname><given-names>B</given-names> </name><name name-style="western"><surname>Hadjitodorov</surname><given-names>S</given-names> </name></person-group><article-title>Acoustic analysis of pathological voices. A voice analysis system for the screening of laryngeal diseases</article-title><source>IEEE Eng Med Biol Mag</source><year>1997</year><volume>16</volume><issue>4</issue><fpage>74</fpage><lpage>82</lpage><pub-id pub-id-type="doi">10.1109/51.603651</pub-id><pub-id pub-id-type="medline">9241523</pub-id></nlm-citation></ref><ref id="ref3"><label>3</label><nlm-citation citation-type="book"><person-group person-group-type="author"><name name-style="western"><surname>Hastie</surname><given-names>T</given-names> </name><name name-style="western"><surname>Tibshirani</surname><given-names>R</given-names> </name><name name-style="western"><surname>Friedman</surname><given-names>J</given-names> </name></person-group><source>The Elements of Statistical Learning: Data Mining, Inference, and Prediction</source><year>2009</year><edition>2</edition><publisher-name>Springer</publisher-name></nlm-citation></ref><ref id="ref4"><label>4</label><nlm-citation citation-type="book"><person-group person-group-type="author"><name name-style="western"><surname>Pearl</surname><given-names>J</given-names> </name></person-group><source>Causality</source><year>2009</year><edition>2</edition><publisher-name>Cambridge University Press</publisher-name></nlm-citation></ref><ref id="ref5"><label>5</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Little</surname><given-names>MA</given-names> </name><name name-style="western"><surname>Varoquaux</surname><given-names>G</given-names> </name><name name-style="western"><surname>Saeb</surname><given-names>S</given-names> </name><etal/></person-group><article-title>Using and understanding cross-validation strategies. Perspectives on Saeb et al</article-title><source>Gigascience</source><year>2017</year><month>05</month><day>1</day><volume>6</volume><issue>5</issue><fpage>1</fpage><lpage>6</lpage><pub-id pub-id-type="doi">10.1093/gigascience/gix020</pub-id><pub-id pub-id-type="medline">28327989</pub-id></nlm-citation></ref><ref id="ref6"><label>6</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Arlot</surname><given-names>S</given-names> </name><name name-style="western"><surname>Celisse</surname><given-names>A</given-names> </name></person-group><article-title>A survey of cross-validation procedures for model selection</article-title><source>Statist Surv</source><year>2010</year><volume>4</volume><fpage>40</fpage><lpage>79</lpage><pub-id pub-id-type="doi">10.1214/09-SS054</pub-id></nlm-citation></ref><ref id="ref7"><label>7</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Omberg</surname><given-names>L</given-names> </name><name name-style="western"><surname>Chaibub Neto</surname><given-names>E</given-names> </name><name name-style="western"><surname>Perumal</surname><given-names>TM</given-names> </name><etal/></person-group><article-title>Remote smartphone monitoring of Parkinson&#x2019;s disease and individual response to therapy</article-title><source>Nat Biotechnol</source><year>2022</year><month>04</month><volume>40</volume><issue>4</issue><fpage>480</fpage><lpage>487</lpage><pub-id pub-id-type="doi">10.1038/s41587-021-00974-9</pub-id><pub-id pub-id-type="medline">34373643</pub-id></nlm-citation></ref><ref id="ref8"><label>8</label><nlm-citation citation-type="book"><person-group person-group-type="author"><name name-style="western"><surname>Aloyayri</surname><given-names>AA</given-names> </name><name name-style="western"><surname>Little</surname><given-names>MA</given-names> </name><name name-style="western"><surname>Zakar</surname><given-names>NA</given-names> </name></person-group><person-group person-group-type="editor"><name name-style="western"><surname>Ni</surname><given-names>H</given-names> </name><name name-style="western"><surname>Cafolla</surname><given-names>D</given-names> </name></person-group><article-title>Causal analysis of Parkinson&#x2019;s motor symptoms using structured smartphone accelerometer data</article-title><source>Artificial Intelligence in Healthcare. AIiH 2026. Lecture Notes in Computer Science, vol 16875</source><year>2027</year><publisher-name>Springer</publisher-name><pub-id pub-id-type="doi">10.1007/978-3-032-35387-0_14</pub-id></nlm-citation></ref><ref id="ref9"><label>9</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Ferrante di Ruffano</surname><given-names>L</given-names> </name><name name-style="western"><surname>Hyde</surname><given-names>CJ</given-names> </name><name name-style="western"><surname>McCaffery</surname><given-names>KJ</given-names> </name><name name-style="western"><surname>Bossuyt</surname><given-names>PMM</given-names> </name><name name-style="western"><surname>Deeks</surname><given-names>JJ</given-names> </name></person-group><article-title>Assessing the value of diagnostic tests: a framework for designing and evaluating trials</article-title><source>BMJ</source><year>2012</year><month>02</month><day>21</day><volume>344</volume><fpage>e686</fpage><pub-id pub-id-type="doi">10.1136/bmj.e686</pub-id><pub-id pub-id-type="medline">22354600</pub-id></nlm-citation></ref></ref-list></back></article>