Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

The leading peer-reviewed journal for digital medicine and health and health care in the internet age. 

Latest Submissions Open for Peer Review

JMIR has been a leader in applying openness, participation, collaboration and other "2.0" ideas to scholarly publishing, and since December 2009 offers open peer review articles, allowing JMIR users to sign themselves up as peer reviewers for specific articles currently considered by the Journal (in addition to author- and editor-selected reviewers).

For a complete list of all submissions across all JMIR journals as well as partner journals, see JMIR Preprints

Note that this is a not a complete list of submissions as authors can opt-out. The list below shows recently submitted articles where submitting authors have not opted-out of open peer-review and where the editor has not made a decision yet. (Note that this feature is for reviewing specific articles - if you just want to sign up as reviewer (and wait for the editor to contact you if articles match your interests), please sign up as reviewer using your profile).

To assign yourself to an article as reviewer, you must have a user account on this site (if you don't have one, register for a free account here) and be logged in (please verify that your email address in your profile is correct).

Add yourself as a peer reviewer to any article by clicking the '+Peer-review Me!+' link under each article. Full instructions on how to complete your review will be sent to you via email shortly after. Do not sign up as peer-reviewer if you have any conflicts of interest (note that we will treat any attempts by authors to sign up as reviewer under a false identity as scientific misconduct and reserve the right to promptly reject the article and inform the host institution).

The standard turnaround time for reviews is currently 2 weeks, and the general aim is to give constructive feedback to the authors and/or to prevent publication of uninteresting or fatally flawed articles. Reviewers will be acknowledged by name if the article is published, but remain anonymous if the article is declined.

The abstracts on this page are unpublished studies - please do not cite them (yet). If you wish to cite them/wish to see them published, write your opinion in the form of a peer-review!

Tip: Include the RSS feed of the JMIR submissions on this page on your homepage, blog, or desktop RSS reader to stay informed about current submissions!

JMIR Submissions under Open Peer Review

↑ Grab this Headline Animator

If you follow us on Twitter, we will also announce new submissions under open peer-review there.

Titles/Abstracts of Articles Currently Open for Review:

  • Validation of a Digital Lifestyle Assessment Tool Against Established Instruments and Its Association with Cognitive Performance

    Background: The Smart Tracker is a brief online questionnaire measuring six lifestyle domains relevant to brain health, including Diet, Exercise, Cognitive Engagement, Social Engagement, Emotional Well-being (sleep and stress/depression), and Physical Well-being. Objective: Our primary objective evaluated criterion validity of the Smart Tracker tool against established instruments measuring similar constructs. We additionally assessed the Smart Tracker’s association with objective cognitive performance in a sample subset who completed the Cogniciti Brain Health Assessment (BHA). Methods: The validation sample included 570 adults (mean age 71.3, SD 9.5 years, age range 22-97; 68.9% [393/570] female) who completed the Smart Tracker alongside five validated instruments: 1) the Eating Pattern Self-Assessment (EPSA; dietary quality), 2) the Short Form International Physical Activity Questionnaire (IPAQ; exercise), 3) the Pittsburgh Sleep Quality Index (PSQI; sleep quality), 4) the Depression Anxiety Stress Scale (DASS-21 short form; psychological wellbeing), and 5) the Cognitive Reserve Index Questionnaire (CRIq; cognitive and social engagement). A nested sub-sample (n=200) additionally completed the BHA. Criterion validity was assessed using Spearman rank correlation coefficients (ρ values) with Benjamini-Hochberg false discovery rate correction to assess the strength and direction of association between each Smart Tracker domain and the corresponding validated instruments. We used a linear regression to assess the association between the Smart Tracker and BHA. Results: The Smart Tracker questionnaire domains showed significant moderate to high correlations with the comparator validated questionnaires assessing similar constructs. Spearman correlations were significantly positive for domains related to Diet (ρ=0.62), Exercise (ρ=0.53), Cognitive and Social Engagement (ρ=0.33), and negative in the expected direction for Emotional Well-Being related to sleep (ρ=-0.65), Stress and Depression (ρ=-0.70; all P<.001). The Smart Tracker Overall score significantly predicted BHA age-normed performance in the cognitive sub-sample (Y = 36.30·X + 47.80; R²=.043; F(1,198)=8.9, P=.003) where each 10-point increase in the Smart Tracker overall score corresponded to approximately 3.6 BHA percentile points. Meaningful ceiling effects were observed for the Smart Tracker Sleep (24.9%) and Stress/Depression (31.8%) sub-scales. Conclusions: The Smart Tracker demonstrates acceptable criterion validity across each of the lifestyle domains it measures and shows a modest but significant association with objective age-adjusted cognitive performance.

  • Managing Uncertainty in AI-Enabled Occupational Injury Claims Review: An Observed-to-Expected Interpretation Framework

    Background: Health and insurance administrators increasingly use statistical models and artificial intelligence methods to identify injury or work categories for closer review. An observed-to-expected (O/E) ratio compares total recorded sickness absence days in a category with the total expected by a prediction model. A ratio greater than 1 indicates more recorded than expected sickness absence. However, it does not explain why the difference occurred or what should be reviewed next. Interpretation may depend on the prediction model, the work information defining the expected comparison, and whether the difference is concentrated among claims with very long durations. Objective: This study aimed to identify occupational injury categories with more recorded sickness absence than expected, compare their O/E results across prediction models and definitions of the expected comparison, and determine whether positive individual claim differences were concentrated among claims with very long durations. Methods: We conducted a retrospective observational study of 86,305 closed occupational injury insurance claims in Hong Kong from 2005 to 2024. Repeated cross fitting with Tweedie generalized linear models estimated expected sickness absence for each claim. Five injury and work category groupings were examined. For each grouping, its defining variable was omitted, so expected duration was based on other recorded characteristics. Categories required an O/E ratio greater than 1.10, a mean recorded minus expected difference of at least 5 days per claim, and at least 500 claims. The selected categories were assessed by replacing the Tweedie model with CatBoost after overall scale alignment, omitting available work information from the expected duration model, and estimating the share of summed positive individual claim differences contributed by claims above the empirical 95th percentile of recorded duration. The first two assessments used a predefined criterion combining relative and absolute changes. Results: Overall recorded and expected sickness absence totals were closely aligned, but 14 categories met all selection criteria. Although CatBoost had lower overall prediction error, only 2 of 14 category comparisons met the predefined change criterion after replacing the Tweedie model with CatBoost. Omitting available work information met the criterion in 6 comparisons. In 10 of 14 selected categories, claims in the upper 5% of recorded durations contributed a greater share of summed positive individual claim differences than the corresponding share in the full cohort. Conclusions: An elevated O/E ratio can identify categories for closer review but cannot explain the observed difference or determine an organizational response. Examining the same selected categories through these three questions separated category selection from subsequent interpretation. These questions provide a transparent way to examine additional evidence before comparisons generated by models are interpreted in organizational review.

  • Background: Psychological counseling chatbots may expand access to mental health support, but evidence has largely focused on scripted systems or general-purpose large language models (LLMs). Whether domain-specific psychological counseling LLMs provide greater clinical benefit remains unclear. Objective: This study evaluated the clinical effectiveness of Nancy, a psychological counseling chatbot built on the domain-specific PsyLLM, relative to ChatGPT, a psychoeducational E-book, and a waitlist control (WLC). Methods: In this 4-arm randomized controlled trial (RCT), 501 Mandarin-speaking adults aged 18 years or older with a Patient Health Questionnaire-9 (PHQ-9) score of 5 or higher or a Generalized Anxiety Disorder-7 (GAD-7) score of 5 or higher were randomized 1:1:1:1 to Nancy, ChatGPT, E-book, or the Waitlist Control (WLC) for 28 days. Primary outcomes were PHQ-9 and GAD-7 scores. Secondary outcomes included affect, perceived empathy, and platform-recorded engagement. Intention-to-treat (ITT) analyses used multiple imputation and baseline-adjusted analysis of covariance; per-protocol (PP) analyses were sensitivity analyses. Results: Nancy showed greater PHQ-9 reductions than ChatGPT (adjusted difference −2.13, 95% CI −3.66 to −0.59; Holm-adjusted p=.007; Cohen d=0.44), E-book (−5.81, 95% CI −7.23 to −4.38; p<.001; d=1.10), and WLC (−2.95, 95% CI −4.35 to −1.54; p<.001; d=0.61). GAD-7 reductions were also greater with Nancy than with ChatGPT (−1.58, 95% CI −3.08 to −0.08; Holm-adjusted p=.04; d=0.35), E-book (−3.35, 95% CI −4.70 to −2.00; p<.001; d=0.74), and WLC (−1.90, 95% CI −3.24 to −0.56; p=.01; d=0.44). Per-protocol findings were consistent. Positive and negative affect did not differ significantly between groups. Perceived empathy was similar between Nancy and ChatGPT (35.49 vs 34.92; p=.69; Hedges g=0.07), whereas conversational intensity was higher with Nancy (median 17.21 vs 16.39 turns per recorded use day; p<.001; Cliff δ=0.29). Conclusions: Nancy produced greater reductions in depressive and anxiety symptoms than ChatGPT, psychoeducation, and waitlist controls over 28 days. These findings suggest that domain-specific therapeutic alignment may provide added clinical value beyond general-purpose conversational support. Clinical Trial: Chinese Clinical Trial Registry (ChiCTR2600121508); https://www.chictr.org.cn/hvshowproject.html?id=298372&v=1.0

  • Integrating Intelligent Communication Robots into Emergency Department Waiting Areas: A Qualitative Exploration of Patient Perceptions

    Background: Emergency departments (EDs) worldwide face increasing challenges due to rising patient volumes, crowding, prolonged waiting times, workforce shortages, and growing administrative demands. In addition to providing urgent medical care, ED staff are required to perform numerous non-clinical tasks, including patient registration, information provision, and redundant communication with patients and accompanying persons about non-medical demands. These challenges can negatively affect patient experience and staff workload. Advances in artificial intelligence (AI) and robotics offer new opportunities to support communication processes and improve operational efficiency in healthcare settings. Objective: This study aimed to evaluate the feasibility, acceptance, and perceived usefulness of an AI-based communication robot deployed in an ED waiting area. Particular attention was given to patients’ prior experiences with technology, contextual factors influencing robot use, social influences on interaction behavior, perceptions of the robotic system, and potential future application scenarios. Methods: A qualitative proof-of-concept study was conducted in the waiting area of the ED at Charité – Universitätsmedizin Berlin. The communication robot TARA (Talking Autonomous Reception Assistant) was implemented alongside a self-check-in terminal to support patient orientation, multilingual communication, and administrative processes. Data were collected during two one-week field phases through participant observations and semi-structured interviews with patients and accompanying persons. The interview guide was informed by the Godspeed Questionnaire Series, covering dimensions such as anthropomorphism, animacy, likeability, perceived intelligence, and safety, as well as by the Unified Theory of Acceptance and Use of Technology (UTAUT) framework, including performance expectancy, effort expectancy, social influence, and facilitating conditions. Data were analyzed using qualitative content analysis according to Mayring. Results: A total of 15 interviews and more than 14 hours of participant observation were conducted. Participants generally viewed the robot positively and highlighted its potential to facilitate orientation, provide information, support registration processes, and overcome language barriers. The robot was often perceived as friendly, approachable, and capable of reducing uncertainty in the waiting area. However, interaction with the robot was strongly influenced by the ED context. Participants prioritized rapid and intuitive registration procedures, while stress, pain, anxiety, and uncertainty regarding the robot’s role frequently limited engagement. Privacy concerns in the open waiting area and a preference for human interaction represented additional barriers. Participants highlighted the need for further development, particularly regarding clearer communication of the robot’s role, more visible benefits for patients, and stronger integration into existing ED workflows, which were considered important prerequisites for future acceptance and use. Conclusions: AI-based communication robots may support patient guidance, multilingual communication, and administrative processes in ED waiting areas while potentially reducing staff workload. Successful implementation depends not only on technical performance but also on intuitive usability, workflow integration, privacy considerations, and the preservation of opportunities for human interaction. Further research is needed to evaluate long-term acceptance and clinical impact in routine emergency care. Clinical Trial: The study is registered with the German Clinical Trials Registry (DRKS-ID DRKS00038333) under https://drks.de/search/en/trial/DRKS00038333.

  • Early versus late internet-based cognitive behavioral therapy after MINOCA or takotsubo syndrome: one-year follow-up of a randomized trial

    Background: Cardiovascular (CV) events such as a myocardial infarction with non-obstructive coronary arteries (MINOCA) or takotsubo syndrome (TS) are associated with heightened levels of stress and anxiety. There are currently no established treatment guidelines for these conditions, but emerging research indicates that internet-based cognitive behavior therapy (ICBT) is effective in the short-term for reducing these symptoms. However, the long-term stability of effects is unknown, as well as the importance of timing of the intervention. Objective: This paper aims to investigate the stability of effects achieved by ICBT after MINOCA or TS for common mental health symptoms, as well as to compare the difference in effect between early and late initiation of the intervention. Methods: This was a parallel, multicenter, randomized controlled trial of ICBT on symptoms of stress and anxiety. The ICBT intervention was a guided, 7–9-week written program, where completion of at least 5/9 steps was considered a minimum dose. Participants were included if they reported elevated stress (Perceived stress scale, 14 items [PSS-14] > 24) or anxiety (Hospital anxiety and depression scale, anxiety subscale [HADS-A] > 7) and were randomized (1:1) to either early intervention group (EIG; immediately after randomization) or late intervention group (LIG; 10-12 weeks after randomization). Mental health outcomes were assessed using PSS-14, HADS, cardiac anxiety questionnaire [CAQ], and impact of event scale [IES]. Outcome data was collected at three follow-up time points: 10-12 (1), 20-22 (2), and 50-52 (3) weeks after randomization. This paper reports results from follow-up 2 and 3. Data was analyzed by the intention-to-treat principle, using linear mixed models, assuming data missing at random. Results: Eighty-eight participants were randomized (EIG: 45; LIG: 43). The response rate at follow-up 2 was 81% and at follow-up 3 was 73%. Eighty-seven percent completed the minimum dose in the EIG, and 51% in the LIG. Intervention effects achieved in the EIG at follow-up 1 remained stable over time. The average trend was lower levels of stress and anxiety in the EIG compared with the LIG, but no differences were statistically significant. Regarding symptoms of post-traumatic stress, the EIG reported significantly lower levels than the LIG, at follow-up 2 (-2.32, 95% CI: -4.42, -0.21, p=0.03). Conclusions: Short-term effects achieved by ICBT after MINOCA or TS for symptoms of stress and anxiety were stable over time and there was little evidence suggesting a difference in effect between early or late initiation of treatment. However, patients in the EIG completed the treatment to a higher degree and reported superior treatment effects on symptoms of post-traumatic stress. These novel findings are important for informing care after a MINOCA or TS, but more research is needed before firm conclusions can be drawn. Clinical Trial: clinicaltrials.gov NCT04178434

  • Comparative Efficacy and Adherence of Digital Health Interventions for Insomnia: A Systematic Review and Network Meta-Analysis

    Background: Insomnia is prevalent and imposes a significant public health burden. Digital health interventions(DHIs) offer convenient and continuous digital management models, substantially improving the accessibility of sleep health management. Nonetheless, the comparative clinical efficacy and adherence across digital interventions with diverse functional architectures remain to be established. Objective: To evaluate the comparative efficacy and adherence of diverse digital health interventions(DHIs) for the treatment of insomnia, and to investigate the associations of delivery support modalities and intervention duration with treatment adherence. Methods: PubMed, Embase, PsycINFO, and the Cochrane Library were systematically searched from inception to July 8,2026. Randomised controlled trials(RCTs) evaluating DHIs for insomnia were eligible. Evidence networks were stratified into human-supported and fully automated sub-networks. Single-arm proportion meta-analysis was used to assess adherence rates, and multivariable meta-regression was used to examine associations with support modality and intervention duration. Risk of bias and certainty of evidence were assessed using RoB 2 and CINeMA, respectively. Results: A total of 138 RCTs involving 24523 participants were included. In the human-supported sub-network, all DHIs except wearable interventions were associated with statistically significant, large reductions in insomnia severity versus passive controls (SMD: -0.79 to -2.71). VR, dBBIs, and AI-driven interventions achieved the highest comparative rankings, whereas App-based CBT-I and Telehealth CBT-I showed comparatively precise and consistent estimates. In the fully automated sub-network, all DHIs except Multimodal CBT-I showed statistically significant, moderate-to-large reductions (SMD: -0.53 to -1.43).VR, wearable interventions, and Web-based CBT-I ranked highest;App-based CBT-I and Web-based CBT-I showed the most consistent evidence. The pooled overall adherence rate was 73.24% (95% CI: 68.48%–77.51%), with human-supported interventions exhibiting significantly higher completion rates than fully automated interventions (80.6% vs. 65.3%). Human support was significantly associated with higher adherence (OR = 2.26, 95% CI: 1.475–3.459). Conversely, each 4-week increase in intervention duration was associated with a 47.7% reduction in the odds of treatment completion (OR = 0.53, 95% CI: 0.34–0.82), with no significant interaction between support mode and duration (OR= 1.01, 95% CI: 0.401–2.561). Conclusions: In human-guided settings, dBBIs,App-based CBT-I,and Telehealth CBT-I, alongside Wearable interventions,App-based CBT-I,and Web-based CBT-I in fully automated settings, represent clinically mature and highly accessible options.Emerging modalities, including VR, wearable interventions,AI-driven interventions, showed high adherence as measured by treatment completion.Human support was associated with higher adherence, whereas longer protocols were associated with lower adherence.

  • The CHASE Study (Comparative Human versus AI Assessment in Tobacco-Dependence Counseling): A Double-Blind Mixed-Methods Evaluation

    Background: Publicly available large language models (LLMs), a class of generative artificial intelligence (AI), are increasingly used by patients seeking health-related advice, while similar systems are being considered for clinical communication. Their ability to provide immediate, personalized support offers opportunities to improve access to health information, but also raises concerns regarding inaccurate or potentially harmful output. Whether high-quality LLM-generated communication is sufficient to support clinical counseling remains unclear. Objective: To address this gap, we compared the communicative quality of LLM-generated and human expert-generated responses to authentic tobacco-dependence counseling questions and examined participant and expert perspectives on LLM-assisted counseling. Methods: This double-blind, mixed-methods study was conducted at the Tobacco Dependence Clinic, LMU University Hospital, Munich. Twenty participant-submitted questions from outpatient smoking-cessation courses and counseling settings were independently answered by three tobacco-dependence experts and a publicly available LLM (Microsoft Copilot). Question submitters and three independent experts rated all responses across five predefined communicative domains (information quality, clarity, safety, practicability, and empathy) using a 5-point Likert scale and completed a blinded source-identification task based on the Turing test. Domain ratings were compared using two-sided paired t-tests. Participants additionally completed questionnaires on attitudes toward LLM-assisted counseling, while experts provided qualitative feedback on the clinical use of LLMs. Results: LLM-generated responses received significantly higher ratings than expert-generated responses across all five communicative domains. In the blinded source-identification task, expert raters correctly identified LLM-generated responses more often than lay participants, though neither group reliably distinguished LLM-generated from human-generated responses. Despite the higher ratings of LLM-generated responses, both participants and experts expressed a preference for human clinicians. Conclusions: Despite higher ratings of LLM-generated responses, participants continued to prefer human clinicians for counseling and experts considered LLMs potentially useful for supportive tasks but not as substitutes for human-led counseling. These findings support a potential role for LLMs as complementary tools in tobacco-dependence care rather than as a stand-alone counseling system. Clinical Trial: The study was retrospectively registered in the German Clinical Trials Register (DRKS-ID; DRKS00035266) on December 12, 2024.

  • Adolescents’ Use of AI Chatbots for Social, Emotional and Mental Health Support: A Cross-Sectional, School-Based Survey

    Background: An ongoing youth mental health crisis has been linked to social technologies, with gendered and developmental differences in youth mental health outcomes related to social technology exposure. Understanding their uses, attitudes about, and experiences of AI can guide policies, clinical practice, and parental and school-based approaches. Objective: Our objective was to characterize adolescents’ uses, attitudes, and experiences regarding GenAI chatbots and examine developmental stage and gender differences. Methods: A STROBE-compliant, cross-sectional evaluation using a youth-informed survey in late 2025 was conducted with school-based adolescents (ages 11-18) across eight sociodemographically diverse campuses. Exploratory survey measures assessed adolescents’ uses, attitudes, and experiences regarding GenAI chatbots. Results: 2,458 adolescents completed the survey (53% of eligible students). Participants included 38% students in grades 6 through 8 and 62% grades 9 through 12; 49% male, 49% female, 0.6% nonbinary/gender fluid; 30% Hispanic/Latine, 21% Black/African American, 18% Asian, 16% White, 9.9% multiracial, and .8% other race. Approximately three-fourths of adolescent participants reported any use of GenAI chatbots, most commonly for socioemotional support (49.5%) and help with school difficulties (62%). The mean age of first use was 12.75 years. Usage was slightly higher among females. Around one-fifth (21%) used AI to share mental health concerns, and around 13% used it to improve their mental health, with both behaviors more commonly reported by females. Early adolescents in grades six to eight (typically ages 11-13) used AI for socioemotional support more often than middle adolescents (typically ages 14-18). Regarding general attitudes, early adolescents held more positive views of AI as a wellness tool than middle adolescents. Additionally, females reported lower general AI optimism and greater AI pessimism than males. Most adolescents (65%) believed AI helped them or a friend, though 9% reported harmful experiences related to AI chatbot use. Conclusions: More adolescents may be using AI chatbots for social, emotional, and mental health support than previously estimated, with first use occurring around age 13. Early adolescence may be an optimal window for AI literacy and mental health literacy interventions, parental conversations about AI use (including risks of sycophancy and dependency), and age-appropriate access restrictions at home and school. Clinical Trial: Not applicable

  • Background: LLMs are increasingly being explored as new tools to support health information seeking. However, stroke survivors may face unique challenges when using digital health technologies due to stroke-related functional impairments. Understanding how stroke survivors interact with LLM systems and adapt AI-generated information to their own health contexts is important for developing more inclusive and patient-centered AI health tools. Objective: To explore how stroke survivors adaptively use large language models (LLMs) for health information seeking and how they understand and apply AI-generated information within their functional, health, and social contexts. Methods: In this qualitative study, 11 stroke survivors were recruited. Participants performed standardized and personalized health information-seeking tasks using three LLMs (DouBao, DeepSeek, and AQ), followed by evaluation of AI-generated responses via the CLEAR scale. Data collection integrated a think-aloud approach with semi-structured interviews, and inductive thematic analysis was applied for data analysis. Results: CLEAR scores indicated generally high evaluations of LLM-generated health information, with researcher ratings slightly higher than participant ratings. Think-aloud findings revealed different processing patterns across tasks: participants primarily verified AI information in the standardized task, whereas the personalized task involved greater contextual adaptation and application, with participants progressively supplementing personal health information. Overall, LLM-based health information seeking among stroke survivors was a process of continuous adaptation. Participants adjusted their interaction approaches according to their functional abilities and progressively supplemented health contexts and clarified their needs during interactions. They also evaluated AI outputs based on credibility cues, uncertainty, personal health conditions, and potential risks. Whether AI-generated recommendations could be translated into actual actions depended on their alignment with patients’ functional abilities and real-life circumstances. Furthermore, this process involved dynamic collaboration among patients, family members, and AI systems. Conclusions: Stroke survivors’ use of large language models for health information seeking is a process of continuous adaptation across interaction, evaluation, and action. Future generative AI systems should tailor interaction modalities and the difficulty of recommendations to patients’ functional abilities, clearly communicate uncertainty and applicability boundaries for high-risk health information, and support family collaboration with patient authorization, thereby helping transform AI-generated health information from being merely accessible to being safely usable and practically actionable.

  • Mechanism Awareness and Unverified Reliance on Generative AI for Health Information in Poland: Preregistered Cross-Sectional Questionnaire Study

    Background: Generative artificial intelligence (AI) chatbots are increasingly used for health questions. A search engine returns a list of sources; a conversational system selects among them and returns one fluent, personalized answer. Detail, coherence, personalization, and agreement with what a user already believes can raise how credible an answer seems without raising how accurate it is. Surveys describe who uses these systems and how far they trust them, but not whether reliance is aligned with what users understand about how the answer was produced. Objective: This study aimed to examine whether understanding of how chatbot answers are produced is associated with reliance on credibility-enhancing conversational features and with self-reported adherence to health advice without independent checking, among adults living in Poland. Methods: This was a preregistered cross-sectional questionnaire study of adults aged 18 years or older living in Poland, recruited nonprobabilistically through an open link and poststratified to Statistics Poland margins for sex, age band, education, settlement size, and region. A mechanism-awareness index (range 0-3) was formed from 3 statements about how answers are produced, with "don't know" scored as not aware. One survey-weighted regression was fitted per registered hypothesis, each adjusted for 7 prespecified covariates including frequency of chatbot use. Confirmatory models were fitted in the subsample reporting health-related chatbot use. Results: Of 3360 submissions, 3348 were valid, of whom 2590 (77.4%) reported health-related chatbot use. The 2 primary awareness associations were approximately null before adjustment and became negative only after conditioning on chatbot-use frequency. Unadjusted estimates were close to null for credibility-cue reliance (β=−0.06, 95% CI −0.16 to 0.03) and for unverified adherence (odds ratio [OR] 0.98, 95% CI 0.81-1.18); after adjustment for the 7 registered covariates, both were negative (β=−0.30, 95% CI −0.38 to −0.22; OR 0.61, 95% CI 0.49-0.76). Replacing general chatbot-use frequency with health-specific use frequency produced a smaller shift (OR 0.81, 95% CI 0.68-0.96). In exploratory analyses, the conditional awareness-adherence association differed across levels of general chatbot use (interaction P<.001). Credibility-cue reliance was strongly associated with epistemic closure—stopping further search after a convincing answer—before and after adjustment (fully adjusted OR 3.75, 95% CI 2.99-4.70) and at every threshold of the outcome (threshold-specific ORs 3.06-9.34). Conclusions: Mechanism knowledge and reliance behavior are not simply aligned. Mechanism awareness showed essentially no marginal association with acting on chatbot advice unchecked, and the registered negative association emerged only once intensity of use was held constant, and not uniformly even then. Reliance on conversational credibility cues, by contrast, was strongly associated with ending further information search. Because the sample overrepresents health-AI users, prevalence estimates are not offered as population parameters. Mechanism education alone should not be treated as a sufficient safety strategy for generative AI in health.

  • Accessibility and Self-Management Outcomes of Patient-Facing Digital Tools for People With Eye Disease and Vision Impairment: Scoping Review

    Background: Digital self-management tools for chronic eye disease are evaluated on device or algorithm performance rather than on whether people already living with functional vision loss can reach and keep using them. Whether accessibility translates into sustained self-management, and whether clinical benefit follows, is largely untested. The role of nursing mediation in that translation is unmapped. Objective: We aimed to map where evidence sits along the pathway from disease severity to clinical outcome for patient-facing digital tools in eye care, and to test whether the apparent absence of nursing-mediation evidence is real or an artefact of database coverage. Methods: Following PRISMA-ScR, we searched PubMed, Web of Science, Scopus and CINAHL Ultimate for English-language records published 2015–2026. Eligibility required eye disease or functional vision impairment, a patient-facing digital tool, and a user-side outcome. Studies were coded on causal-chain position, tool class, vision status, outcome type and evidence grade. Results: Of 7,536 records, 97 studies were included. Evidence concentrated in one link: 85 studies (87.6%) addressed interface accessibility → self-management behaviour, against 9 and 3 for the adjacent links; 76 stopped at an accessibility or usability endpoint and only 6 reached a clinical outcome. Only 21 (22%) measured any causal edge. The nursing-mediation layer is a genuine absence within our criteria — no study treated nursing attribution as an examined variable. A severity gradient runs in the same direction but is not significant (Fisher P=.12). Conclusions: Research has documented whether users can operate these tools, leaving untested whether operation changes outcomes or whether nursing mediation makes it possible. We term this the accessibility paradox and specify four measurable breakpoints in the assess–adapt–teach–follow-up sequence, so that the zero finding is refutable by design.

  • Generative AI and Medical Manuscript Integrity

    Abstract Generative artificial intelligence is increasingly used in medical manuscript preparation. Large language models may improve readability, organization, and access to scientific communication, particularly for authors who are not native English speakers. The same technologies may also create new vulnerabilities for research integrity by generating fabricated or misused references, persuasive but unsupported interpretation, and polished manuscripts that conceal weak evidence or inadequate human intellectual contribution. This position paper analyzes generative AI in medical publishing as a challenge for artificial intelligence in medicine, because unreliable manuscripts may enter evidence synthesis, clinical guidelines, educational materials, and ultimately clinical decision making. We propose a three-group framework: honest researchers who use AI to improve communication; dishonest researchers who use AI to accelerate deception; and an intermediate group of fundamentally honest authors who may adopt shortcuts under academic or financial pressure. The central risk is not visibly artificial text but credible-looking manuscripts whose form imitates scholarship while their evidentiary substance remains weak or unverified. We recommend transparent AI disclosure, manual reference verification, data availability, reviewer education, editorial screening for integrity risks, and reform of promotion systems that reward publication volume over scientific quality.

  • Human Connectedness and the Limits of Artificial Intelligence in School Counselling for Gay Learners in South Africa: A Qualitative Study

    Background: Conversational artificial intelligence (AI) is now an ordinary feature of adolescent help seeking, and users reportedly form alliance-like bonds with such systems. Debate about AI in mental health care has been dominated by capability benchmarking, and has been conducted almost entirely with reference to well-resourced Global North settings. Little is known about what these systems would have to substitute for in low-resource school settings, or in work with sexual minority learners, for whom disclosure to an dules in a school carries distinctive risks. Objective: This study examined how school-based counselling practitioners in South Africa describe their work with gay and same-sex attracted learners at risk of suicide, in order to specify the constituent properties of relational support, determine which are plausibly substitutable by current or foreseeable AI systems , and established where AI may and may not be legitimately deployed. Methods: We conducted a qualitative descriptive study using semi-structured individual interviews with 11 school-based counselling practitioners working in three South African provinces, purposively sampled on current practice and experience supporting learners who had disclosed same-sex attraction and suicidal ideation. Data were analysed using reflexive thematic analysis; reporting follows COREQ-32. Participants were not asked about artificial intelligence and none raised it, the comparison between their accounts and the AI capabilities is an analytic overlay contributed by the authors and declared throughout. Results: Five constituents of relational support were identified: trust as an accumulated and accountable achievement rather than a procedural formality; the reading of distress in embodied and paralinguistic cues; identify affirmations as a staged developmental process rather than a single affirming response; advocacy conducted beyond the counselling room within school and family systems; and the moral and emotional cost borne by the practitioner. A sixth, contextual finding cut across the others: participants described schools in which teaching and support staff were themselves sources of stigma and bullying, and in which no counselling resource existed at all. Mapping these against the AI literature yielded a three-part taxonomy of limitation: contingent-technical, which current systems are already overcoming; institutional-structural, durable but contingent on law and regulation; and constitutive, concerning reciprocal vulnerability and moral answerability, which do not reduce to technical capability. Conclusions: Arguments for the irreplaceability of human practitioners frequently rest on contingent-technical claims that are already obsolete, weakening an otherwise sound position. The durable case rests on institutional standing and on constitutive reciprocity. We propose a permissibility framework specifying where AI may operate independently, where it may operate only as an adjunct, and where it should not operate at all in school-based support for sexual minority learners, together with the safeguards each tier requires. We further argue that the common recommendation of AI as a bridge to human care presupposes that the human destination is safe, an assumption our data do not support in all settings. South Africa currently has no regulatory instrument governing AI-delivered psychological services; we set out what such an instrument would need to address

  • Applications, Predictive Models, and Clinical Integration of Artificial Intelligence in Autologous Breast Reconstruction: Scoping Review

    Background: Artificial intelligence has undergone rapid development in recent years and is increasingly integrated into various medical specialties, including plastic and reconstructive surgery. However, a comprehensive mapping of AI applications specific to autologous breast reconstruction is required. Objective: This scoping review aims to systematically map the current literature regarding artificial intelligence applications in autologous breast reconstruction. Methods: A scoping review was conducted using the MEDLINE, Scopus and Embase databases. Full-text original articles investigating the use of AI in autologous breast reconstruction published between January 1, 2020, to June 30, 2026, were included. Abstracts, case reports, animal studies and studies evaluating non-autologous breast reconstruction exclusively were excluded. Thirty studies met the final eligibility criteria. Results: Among the 30 included studies, AI was predominantly applied in the preoperative phase, with emerging applications in intraoperative and postoperative care. Preoperatively, AI was utilized for risk and complication prediction, patient-reported outcomes, patient satisfaction modeling, automated imaging analysis, and interactive patient counseling. AI was applied intraoperatively to answer clinical queries and quantify perforators in DIEP flap surgery. Postoperatively, deep learning models were applied to postoperative assessment and patient-centered outcome evaluations. Conclusions: AI is currently applied extensively in the preoperative phase of autologous breast reconstruction for predictive modelling, imaging automation and patient counseling. Intraoperative and postoperative applications are equally emerging. Further studies are required to validate these models on larger, diverse datasets to ensure safety and accuracy before routine clinical adoption.

  • From Computational Capability to Healthcare Value: Capability–Value Coupling as an Explanatory Focus for Healthcare AI Translation

    Healthcare artificial intelligence (AI) is advancing rapidly, yet improvements in computational capability do not reliably translate into corresponding improvements in healthcare. Existing implementation, sociotechnical, realist, clinical utility, and AI evaluation approaches provide substantial resources for understanding how AI is evaluated, adopted, embedded, and made consequential in healthcare systems. However, evidence about changes in computational capability and changes in realised healthcare value is often generated across different research traditions, stages of the technology lifecycle, and units of analysis, leaving their relationship insufficiently examined as an explicit object of explanation. We describe this relationship as capability–value coupling: the extent and manner in which variation in task-relevant computational capability contributes to variation in realised healthcare value through context-dependent mechanisms over time. We argue that studying this relationship requires investigators to specify the capability expected to matter, determine whether additional capability creates actionable utility, examine the processes through which it may influence healthcare, prespecify the relevant healthcare value, and distinguish co-occurring capability–value alignment from evidence supporting a coupling explanation. Longitudinal and comparative investigation can further identify mechanisms and boundary conditions under which additional capability becomes consequential, is amplified, attenuated, delayed, or fails to generate additional value. Capability–value coupling is proposed not as a new translational framework, but as an explanatory focus for connecting existing forms of evidence and developing more cumulative knowledge about when, how, and why advances in AI become advances in healthcare.

  • Opioid equipotency conversion with large language models: A controlled simulation study to evaluate clinical safety.

    Background: Large language models (LLMs) are increasingly used in medical education and clinical decision-making, but their reliability in high-risk medication dosing remains unclear. Opioid rotation is a common task requiring precise calculations where errors may result in overdose or inadequate pain relief. Objective: The objectives of this study are to assess the performance of commercial LLMS in opioid conversions and further analyze types of errors in process. Methods: Thirteen LLMs were tested using an API-based framework to ensure independent queries across trials. First, fictional clinical scenarios were tested to simulate real-world clinical situations involving opioid rotation; to test the effects of changes in wording, scenarios were revised into 4 “vignettes” showing the same clinical situation. Next, opioid pairs were tested with a random-dose paradigm across a clinically-pertinent range (5-120 mg daily morphine equivalents). LLM outputs were compared with expected values derived from reference standards. Accuracy was assessed using predefined safety thresholds: tight accuracy (0.85–1.15x expected dose) and broad accuracy (0.6–1.7x). We tested models naively and with prompts augmented with reference tables and unit explanations. Results: Naive models generally exhibited low tight-range accuracy across opioid pairs. For any given opioid pair, each model would consistently produce similar incorrect conversion ratios despite wide variability across opioid pairs and language models. Vignette wording changes accounted for 76% of within-scenario response variance. Reference-based prompt augmentation significantly improved performance, with over half of models achieving high proportions of conversions within tight accuracy for morphine-equivalent conversions. Conclusions: While commercial LLMs demonstrated variable accuracy in the native state, prompt augmentation significantly improved their performance.

  • Patient Experiences with Generative AI in Healthcare: A Systematic Literature Review and Development of a Conceptual Framework

    Background: Previous research on patient experiences with generative artificial intelligence (GenAI) has primarily relied on static technology acceptance models to evaluate the use of AI tools in healthcare. However, it has been poorly studied which mechanisms shape acceptance, empowerment, resistance, and the changing relationship with healthcare professionals. Objective: This Systematic Literature Review (SLR) evaluates patient experience with GenAI in healthcare, thereby aiming to understand the patient– GenAI interaction in healthcare and frame the dynamics in patient experience. Methods: The review consists of empirical studies published in 2022–2026 across six major databases and records. It is conducted following the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA-2020) guidelines. It synthesised diverse methodologies and tools. Records were screened using Rayyan. Researchers assessed study quality using the Mixed Methods Appraisal Tool (MMAT), evaluated theme confidence using the Confidence in the Evidence from Reviews of Qualitative Research (GRADE-CERQUAL) framework, coded data with MAXQDA, and synthesised data using the Thomas and Harden (2008) framework. Results: Key findings suggest differing patterns of Patient – GenAI interactions according to their control over GenAI, the authority the GenAI assumes, experience, agency and literacy of the patient. Across studies, prior experience and health/AI literacy were associated with differences in perceptions of empathy, privacy, trust, and clinical autonomy, although the direction and strength differed by context. Furthermore, patients appear to value conversational anonymity while often having limited awareness of data handling. As GenAI tools transition from administrative tasks to clinical operations, some patients become more selective or resistant to the autonomous clinical roles, favouring the human-to-human diagnostic relation. The review proposes the Dual-Pathway Patient-GenAI Conceptual Model (DP-PGCM) to map patient-controlled and provider-deployed/controlled interaction trajectories. Conclusions: Patients’ GenAI use is a dynamic process. It is role-dependent rather than technology-wide. Patient experience appears to be structured by who controls the tool/system, what clinical authority the tool is granted, and how patients interpret its empathy, privacy, and epistemic reliability. Clinical Trial: PROSPERO registration number: (CRD420261362444). The title and keywords are revised to include the conceptual model.

  • Exploring the Application of Robotics in Cardiac Rehabilitation: Scoping Review

    Background: Cardiac rehabilitation is a core measure for secondary prevention of cardiovascular disease, but global participation rates remain consistently low. Robotics, with its precise movement assistance capabilities and personalized interactive functions, offers a new approach to overcoming the limitations of traditional rehabilitation models. However, research in this field is diverse in design and evidence remains scattered, and there is a lack of systematic review and integration. Objective: This study aims to use a scoping review methodology to systematically review the current status of robotics applications in cardiac rehabilitation for patients with cardiovascular disease. It seeks to summarize the types of technologies, target populations, rehabilitation stages, and outcome measures; clarify the scope and distribution characteristics of existing evidence; identify knowledge gaps; and provide a reference basis for future research design and clinical practice. Methods: We conducted this review from March to July 2026. With one lead researcher and two professional librarians, we developed search strategies and systematically searched three Chinese databases (CNKI, Wanfang, VIP) and four English databases (PubMed, Web of Science, Cochrane Library, Embase) from inception to April 4, 2026. Two independent reviewers screened studies, extracted key data, and resolved disagreements by consulting two additional reviewers. Results: This study ultimately included 10 articles involving 587 patients with cardiovascular disease from 5 countries. The published articles spanned the years 2015 to 2025, and the study designs included 3 randomized controlled trials and 7 non-randomized studies. Robotic technologies can be categorized into two main types: physical rehabilitation robots—which encompass four technical architectures, including exoskeleton gait training, balance training, motorized resistance training, and soft wearable devices, accounted for 70% of the included studies; social assistance robots, with voice interaction and personalized feedback as their core functions, accounted for 30%. In terms of rehabilitation phases, 7 studies focused on Phase II outpatient rehabilitation, 3 involved Phase I inpatient rehabilitation, and there were no studies on the Phase III community maintenance phase. Outcome measures covered four dimensions: physiological, functional, psychosocial, and safety. Robot-assisted training demonstrated positive effects in improving cardiopulmonary function, enhancing motor abilities, and increasing patient engagement, and its safety was preliminarily validated across populations with varying disease severities. Conclusions: Robotic technology is driving the transformation of cardiac rehabilitation from a traditional model dominated by human supervision to an intelligent model of human-robot collaboration. The two types of robots complement each other functionally, collectively addressing the core challenges of traditional rehabilitation regarding individualized supervision and maintaining motivation among patients with cardiovascular disease. This review provides a conceptual framework for future research and helps advance the field from fragmented exploration toward systematic and standardized development.

  • Evaluative Stance Toward Artificial Intelligence in the Medical Literature: Infoveillance Study of 16,749 High-Quartile Journal Abstracts, 2021-2026

    Background: Medical artificial intelligence (AI) publications report performance while also framing AI as beneficial, uncertain, or risky, which has not been measured at scale. The literature is itself an information environment clinicians rely on, yet health AI sentiment has been measured mainly in public online discourse, not in medical journals. Objective: We aimed to measure the evaluative stance toward AI in high-quartile medical journal abstracts from January 2021 to April 2026, and to characterize it across time, concern themes, failure mechanisms, specialties, first-author country, and publication format. Methods: We conducted a large language model (LLM)-assisted computational content analysis of PubMed records from top-two-quartile (Q1/Q2) SCImago journals. Of 97,492 eligible records, an AI-engagement prefilter retained 16,759 discourse or evaluative records, and stance classification yielded 16,749 valid labels on an ordered scale of Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Critical stance was Alarm plus Caution; favourable stance was Cautious Optimism plus Advocacy. Critical records were further classified for concern theme and, on two axes, for failure mechanism, which distinguished fabrication from factual error, and model type. Every LLM step was validated against blinded human coding (prefilter Cohen kappa 0.51; stance quadratic-weighted kappa 0.79, 95% CI 0.72-0.84; specialty kappa 0.75; model type kappa 0.84; failure mode kappa 0.57). Results: Cautious Optimism was the majority stance (62.8%) and 30.8% of records were critical. Advocacy fell monotonically from 2.9% (2021) to 0.6% (partial 2026; Spearman rho -1.00), while critical share rose modestly from 25.4% (95% CI 22.9-28.1) to 32.6% (30.9-34.4), rising within both prefilter genres. Among critical abstracts, patient safety rose 11.0 percentage points and errors 30.9, whereas regulation fell 22.0 in prevalence share. Within the hallucination-gated subset (n=2226), factual error was the majority mechanism every year (56% to 90%), while fabrication-involved papers reached about 21% to 23% from 2023 onward; fabrication estimates are an upper bound. Critical rate varied roughly threefold across 18 specialty and domain categories (51.3% in Mental Health and Psychiatry to 16.8% in Cardiology), was inversely associated with US Food and Drug Administration cleared-device availability (rho -0.65, two-sided p=.004), and varied by first-author country (34.5% United States, 16.2% China). Reviews were the least critical (24.9%) and most favourable (73.1%) format. Conclusions: Medical-AI abstracts moved away from unqualified promotion toward more qualified assessment, not toward broad opposition. Two implications follow for how the literature is written and read. That reviews were the least critical and most favourable format may matter for readers, guideline developers, and educators. Critical reporting should distinguish factual error from fabrication rather than treating hallucination as one category, which bears on how journals, reviewers, and reporting standards describe language model failure. Throughout our investigation, we measure published discourse, not AI capability, nor whether these stances are correct.

  • Background: Selfies combine several appearance-related activities into a single routine: taking photographs, choosing among them, editing selected images, and deciding whether to post. Research has linked some of these activities, particularly comparison and editing, to poorer body image, while posting and supportive feedback can have different associations. Less is known about how the reasons people use a platform relate to this routine, or where body appreciation fits within those relationships on Chinese social media. Objective: This study examined the direct and indirect associations between four UGT-derived motivations for social media use (information seeking, social interaction, entertainment, and status seeking), body appreciation, and selfie engagement among young Chinese women who use Rednote (Xiaohongshu). Methods: A cross-sectional, self-administered web-based survey was distributed via Rednote and WeChat and hosted on Wenjuanxing in June 2026, following a May 2026 pilot (n=72). Eligible respondents were Chinese women aged 18 to 35 years who were active Rednote users. Of 567 submissions, 21 failed eligibility screening, and 94 were classified as straight-lined in the study's cleaning record, leaving 452 analyzable responses. Motivations were measured with a 12-item, 4-dimension UGT scale; body appreciation with the Body Appreciation Scale-2; and selfie engagement with the 15-item Selfie Engagement Scale, modelled as a reflective-formative higher-order construct comprising preoccupation, selection, editing, and posting. Partial least squares structural equation modelling used 5000 bootstrap resamples. Results: The measurement model met the prespecified criteria (Cronbach αlpha, 0.707-0.917; average variance extracted 0.574-0.697; all heterotrait-monotrait ratios <0.90). All four motivations were positively associated with body appreciation, jointly explaining 65.4% of its variance: social interaction (β=.310, P<.001), information seeking (β=.300, P<.001), entertainment (β=.233, P<.001), and status seeking (β=.124, P=.01). Three were directly associated with selfie engagement: information seeking (β=.222, P<.001), status seeking (β=.221, P<.001), and social interaction (β=.139, P=.02). Entertainment was not (β=.068, P=.07). Body appreciation had the largest standardized coefficient in the selfie engagement equation (β=.291, P<.001; R²=0.623). All four statistical indirect effects through body appreciation were significant (β=.036-.090; all bias-corrected 95% CIs excluded zero). Entertainment showed an indirect-only association; the other three were complementary Conclusions: In the fitted model, body appreciation accounted for a modest part of the association between each motivation and selfie engagement. Entertainment showed an indirect association only. The positive association between status seeking and body appreciation resembles an earlier finding from Korea, but culture was not measured and should not be treated as the explanation. The results support closer study of positive and negative body image processes together, using longitudinal data and measures that distinguish selfie selection, editing, and posting Clinical Trial: Not Applicable

  • Evaluating AI-Generated Physical Activity Recommendations in Primary Care: Qualitative Framework Development Study

    Background: In primary care, artificial intelligence (AI) is increasingly being developed to support personalised physical activity recommendations for people living with long-term conditions by integrating clinical, behavioural and contextual information. As integration occurs, patients, carers and healthcare professionals must decide whether AI-generated recommendations should influence clinical care. Existing qualitative studies describe stakeholder perspectives but provide limited explanation of the reasoning processes underpinning these decisions. Objective: This study aimed to develop a stakeholder-informed conceptual framework describing how people living with long-term conditions, carers and healthcare professionals evaluate AI-generated physical activity recommendations in primary care. Methods: Framework analysis was used to explore how stakeholders evaluated AI-generated physical activity recommendations. We conducted a secondary analysis of transcripts from four focus groups, which were inductively coded to identify participants’ perceptions of AI-generated recommendations. Through iterative researcher–Large Language Model dialogue, codes were synthesised into higher-order concepts across stakeholder groups. These concepts were subsequently developed into an explanatory framework describing the reasoning processes underpinning stakeholders’ evaluation of AI-generated physical activity recommendations in primary care. Finally, concepts from the COM-B model were mapped onto the framework to examine how capability, opportunity and motivation may influence responses to, and engagement with, AI-generated physical activity recommendations. Results: Stakeholders evaluated AI-generated physical activity recommendations through four interacting constructs. Decisions depended on confidence in the recommendation (construct one = credibility), its alignment with individual health and circumstances (construct two = personal fit), appropriate future clinician involvement in sustaining recommended physical activity (construct three = supported use), with any behaviour change underpinned by the individual’s capacity, motivation and opportunity to act on recommendations (construct four = behavioural appraisal). Rather than relying on a single consideration, participants integrated these constructs before deciding whether AI-generated recommendations should influence care. Conclusions: The stakeholder-informed conceptual framework developed as part of this study demonstrates that successful implementation of AI-enabled physical activity recommendations depends on how recommendations are evaluated within clinical decision-making, rather than on algorithmic performance alone. By identifying the reasoning processes used by patients, carers and healthcare professionals, and considering personal behaviour change barriers and facilitators, the framework provides practical guidance for designing AI systems that support shared decision-making, facilitate clinical implementation and improve the acceptability of AI-generated recommendations in primary care. The study also demonstrates the potential of researcher-directed LLM-assisted secondary qualitative analysis to support transparent conceptual framework development while maintaining researcher oversight.

  • Large Language Models' Causal Reasoning in Epidemiology: Development and Validation of the CausalNHANES Automated Benchmark

    Background: Evaluating Large Language Models (LLMs) on epidemiological causal inference tasks traditionally relies on expensive human expert scoring, which suffers from high cost, poor reproducibility, and scalability bottlenecks. Automated benchmarks are urgently needed to rigorously assess whether LLMs truly understand causal inference principles — the cornerstone of epidemiological methodology. Objective: We aimed to develop and validate CausalNHANES, a fully automated benchmark that evaluates LLMs' causal reasoning capabilities using the National Health and Nutrition Examination Survey (NHANES) dataset as the empirical substrate, without requiring human expert scoring. Methods: We constructed a 3+1 evaluation framework (three core causal reasoning tasks plus one safety gate) comprising (1) directed acyclic graph construction, (2) confounder identification, (3) causal inference from descriptive statistics, and (4) counterfactual reasoning. Evaluation relied on two objective computational anchors: literature consensus anchoring via PubMed-based retrieval and logical rigidity anchoring via formal causal logic rule verification. Cross-model consensus was computed post hoc as a supplementary qualitative probe for error-pattern clustering, not as a quantitative scoring component. We systematically evaluated 9 LLMs across 80 test scenarios (60 primary tasks plus 20 negative-controls). The confounder identification layer operated as a binary pass/fail sanity check and was excluded from the composite score. Innovation was redefined as strategy finesse multiplied by a logic-rigidity admission term (squared penalty for violations) and negative-control pass rate. The negative-control scenarios served as a prerequisite disqualification filter: any model hallucinating a causal relationship in a negative-control scenario was excluded from the main leaderboards and listed separately in a disqualified cohort. Because every evaluated model failed at least one negative-control, no model qualified for the ordinal tier leaderboards. This universal failure became the primary analytical focus. Results: The formal experiment evaluated 9 state-of-the-art LLMs across 720 scenario-model pairs (180 negative-control scenarios). Strikingly, all 9 models failed at least one negative-control scenario, yielding a zero-percent qualification rate for the main leaderboards. No model achieved Platinum tier (negative-control pass plus zero logical violations). ERNIE-4.5 and Doubao approached Gold-tier eligibility (negative-control pass with low violation rates) but were ultimately disqualified by hallucinated causal claims in negative-control tasks. Kimi-K2.6 exhibited the highest logical violation rate (49.0%). The pattern was consistent across model families and geographic origins: Chinese-origin and Western-origin models were equally susceptible to hallucinating causal relationships where none exist. Notably, Western flagship models (GPT-4o, Claude-Sonnet-5, Gemini-3.6-Flash) also failed negative-controls, confirming that the vulnerability is architectural rather than geographically or commercially bounded. Conclusions: CausalNHANES reveals a sobering finding: despite rapid advances in general reasoning, current state-of-the-art LLMs are not yet reliable for epidemiological causal inference. The universal failure on negative-controls—a task requiring only the ability to state 'no causal relationship'—suggests that models lack genuine causal reasoning capability and instead rely on pattern-matching heuristics that overfit to epidemiological vocabulary. This failure-mode analysis, rather than a performance ranking, constitutes the primary scientific contribution. The framework itself remains fully reproducible, cost-efficient, and scalable, providing a methodological paradigm for safety-first AI evaluation in high-stakes epidemiological and digital health domains.

  • Digital Engagement among Culturally and Linguistically Diverse Population: A qualitative study

    Background: Digital engagement among adults can impact healthcare outcomes and patient engagement; however there are barriers to utilizing digital tools. Objective: This study explored those barriers and facilitators to digital engagement among English and Spanish speaking adults in Houston, Texas. Methods: We conducted 4 focus groups, 2 in English and 2 in Spanish. Results: In this paper authors present the themes that emerged among them including, (1) Digital Literacy and Technology, (2) Trust, Privacy, and Security Concerns, (3) Support System and Workarounds, (4) Usability and Accessibility, (5) Communication Preferences, and (6) Digital Health Engagement. Results suggest participants desire simplicity, quick access, and easy to understand digital tools. Participants expressed familiarity with utilizing digital tools and accepted how advanced technology was a part of the world today. Participants shared frustrations and fears when providing health data to their providers. They also provided insights into educational information that might help facilitate utilization of digital tools. Conclusions: Findings from our study could inform the importance of customizing and personalizing digital tools for culturally and linguistically diverse populations served within hospital systems to optimize their adaptation and effectiveness.

  • The Digital Health literacy of Adolescents: a Scoping Review

    Background: Social media and the internet are increasingly popular sources of health information for adolescents. While access to vast amounts of information in the digital space provides opportunities for empowerment and learning, the lack of regulation and widespread misinformation pose risks, which may potentially impact adolescents’ health behaviours and outcomes. Therefore, the ability to understand and critically evaluate the trustworthiness of the information accessed, key components of digital health literacy, are becoming increasingly important skills. Objective: To assess the extent and breadth of the literature available investigating the digital health literacy of adolescents, reporting how it has been defined and the measurement tools and approaches used. Methods: This scoping review is guided by Arksey and O’Malley’s methodology for scoping reviews, with further guidance from Levac et al and the Joanna Briggs Institute’s guidance for scoping reviews. It is reported using PRISMA guidelines. Searches for relevant studies were conducted across five electronic databases: PubMed, CINAHL, Web of Science, Education Source and Scopus. Database searching was supplanted with citation searching. Empirical studies that measure the digital health literacy of adolescents were eligible for inclusion. Results: Fifty-eight articles were identified for inclusion, published between 2011 and 2025. The concept of digital health literacy is evolving with the expansion of the digital space. No adolescent specific definition or framework was identified in the literature. Most studies utilised self-report measurement instruments. Five validated self-report measurement tools were identified and six studies employed objective performance-based measurement approaches. Three studies used a combination of a self-report measure and objective measurement. Conclusions: The research is expanding rapidly; however, it lacks conceptual clarity with no consensus on a single approach of measurement. Existing measurement tools require greater flexibility and ongoing revision to reflect the speed at which digitalisation and technology are evolving. Adolescent involvement in future development of measurement instruments may enhance their relevance for this demographic.

  • A Rapid, Reproducible AI-Assisted Workflow for Generating Course-Aligned Anki Decks in Biomedical Sciences Education: A Tutorial

    • This manuscript needs more reviewers

    Background: The Marian University Biomedical Science (BMS) Master’s program is an intensive graduate-level program designed to mimic the rigor of the first year of medical school. This requires students to implement self-directed strategies to manage cognitive load and support retention. Anki is a digital flashcard program that enables spaced repetition and active recall, but manual flashcard generation is time consuming. Here we describe a reproducible workflow using ChatGPT 5.2 to generate Anki decks aligned to course objectives and lecture content. Objective: To describe and demonstrate a reproducible AI-assisted workflow for rapidly generating course-aligned Anki decks from instructor-provided lecture materials in medical sciences education curriculum. Methods: Lecture slides were downloaded as editable PowerPoint files and modified during class to remove identifying information, delete non-testable information, clarify key points, and condense extra content. After lecture, edited slides were exported as a PDF and provided to a ChatGPT 5.2 Thinking session with a lecture-specific prompt. ChatGPT generated two import-ready CSV files per lecture, which were imported into Anki. Results: Across six completed examination blocks, 17,222 total Anki cards were generated from BMS coursework. Mean deck size was 106 cards per lecture. Cards comprised 56% Basic (9,644) and 44% Cloze (7,578). Deck generation required ~3-5 minutes per lecture, and >99% of cards were retained in the decks with only minor edits for clarity or focus Conclusions: This AI-assisted workflow functions as a scalable learner-support tool to rapidly produce high-yield, objective-aligned Anki decks in an accelerated graduate curriculum, enabling immediate post-lecture studying and recall practice customizable to course content and individual student needs. To our knowledge, this represents one of the largest implementations of AI-assisted flashcard generation within medical curriculum.

  • AI-Based Conversational Agents Show Potential to Improve Hypertension and Type 2 Diabetes Management: A Systematic Review

    Background: Artificial intelligence (AI)–based conversational agents (CAs), including chatbots and virtual assistants, are increasingly being used to support chronic disease self-management. However, evidence regarding their implementation, effectiveness, and equity implications in management of hypertension and type 2 diabetes has not been synthesized. Objective: To synthesize evidence on the feasibility, acceptability, implementation, effectiveness, and equity of AI-based conversational agents for adults with hypertension and/or type 2 diabetes. Methods: We conducted a systematic review in accordance with PRISMA 2020 guidelines. EMBASE, Web of Science, MEDLINE, and CINAHL were searched for studies published between 2015 and 2025. Google Scholar and reference lists of relevant reviews and included studies were searched to identify additional records. Eligible studies evaluated AI-based conversational agents among adults with hypertension and/or type 2 diabetes and reported clinical, behavioral, patient-centered, or implementation-related outcomes. Study quality was assessed using the Mixed Methods Appraisal Tool, and findings were synthesized narratively following the Synthesis Without Meta-analysis guidelines. The certainty of evidence was assessed using the GRADE approach. Results: Sixteen studies met the inclusion criteria, encompassing diverse study designs and settings across 4 low-/middle-income and 11 high-income countries. Sample sizes ranged from 10 to 64,679 participants, primarily including adults with type 2 diabetes, hypertension, obesity, or cardiometabolic conditions. AI-based conversational agents were delivered through mobile applications, SMS/text messaging, and other messaging platforms. Interventions commonly supported medication adherence, blood pressure and glucose monitoring, lifestyle modification, psychosocial support, and self-management education. Statistically significant improvements were reported in glycemic outcomes, with HbA1c reductions ranging from 0.30% to 1.04%. Blood pressure benefits were observed primarily for systolic blood pressure and overall blood pressure control, with systolic blood pressure reductions ranging from 3.0 to 6.5 mmHg. Behavioral and patient-centered outcomes also improved, including medication adherence (60.0% to 82.2%), depression symptoms, diabetes distress, and health-related quality of life. Implementation findings indicated high feasibility and acceptability, particularly when conversational agents were integrated into structured care pathways involving nurse- or pharmacist-led support. Equity-related findings indicated that CAs were implemented among low-income, uninsured, older, and ethnically diverse populations, but digital access and literacy remained barriers; equity-promoting strategies included SMS delivery, multilingual content, and subsidized devices. Declining engagement over time was also reported. Overall, certainty of evidence ranged from very low to low due to methodological limitations, heterogeneity, incomplete outcome data, and potential publication bias. Conclusions: AI-based conversational agents show promise for improving self-management, behavioral outcomes, and selected clinical outcomes among adults with hypertension and type 2 diabetes. Their effectiveness may be greater when integrated into broader healthcare systems and sustained models of care. Addressing digital access, literacy, and representation will be essential to ensure equitable implementation and benefit. More rigorous, large-scale studies with longer follow-up are needed to evaluate clinical effectiveness, sustainability, implementation, and equitable access across diverse populations.

  • Background: Physicians caring for patients with complex care needs often need to review large volumes of fragmented clinical information from several sources. Whether better access to more clinical notes is sufficient, or whether information must also be summarised and visualised to support clinical work, remains uncertain. Objective: The aim was to determine whether a digital tool integrating clinical notes from several sources with elements that summarise and visualise information (“DigiTeam”) improves the quality of physicians’ written case summaries and care plans for complex patient cases, and whether any improvement reflects summarising and visualising information, access to several sources, or both. Methods: A three-arm, parallel-group, superiority randomised controlled trial among physicians in Norway. Participants were randomised to one of three digital versions: access to clinical notes from a single-source, access to clinical notes from several sources, or DigiTeam, which combined notes from several sources with summaries and visualisations. Participants reviewed three real-world pseudonymised patient cases with complex care needs and wrote a patient summary and care plan for each case. The primary outcome was the quality of the written summaries and care plans, assessed against predefined expert-derived criteria. Results: Of 93 randomised physicians, 90 contributed data. Access to several sources did not improve overall quality compared with single-source access (mean difference 0.5, 95% CI −2.3 to 3.3; p=0.970). DigiTeam produced higher overall quality scores than multi-sources access without summarising and visualising elements (mean difference 6.49, 95% CI 3.2 to 9.8; p<0.001) and than single-source access (mean difference 6.95, 95% CI 3.7 to 10.2; p<0.001). DigiTeam also improved perceived information support, while time spent and usability did not differ between groups. Conclusions: A digital tool that summarised and visualised clinical information improved the quality of physicians’ summaries and care plans without increasing time use or reducing usability. In contrast, access to more clinical notes alone did not improve quality. These findings suggest that, for complex patient cases, how information is presented matters more than how much information is available. In short, summarised and visualised information improved quality; more information alone did not. Clinical Trial: ClinicalTrials.gov, NCT07142616, registered on 21 August 2025

  • Background: Perioperative health education is important for postoperative recovery among patients with gallstones. However, traditional health education faces challenges related to insufficient continuity and personalization during perioperative information delivery. Objective: This study aimed to apply an AI agent–based health education system in perioperative health education for patients with gallstones and evaluate its effects on improving postoperative recovery outcomes. Methods: This study conducted a quasi-experimental research on patients undergoing laparoscopic cholecystectomy in the hepatobiliary surgery department of a tertiary grade A hospital in China. Patients were grouped according to their admission sequence. Patients in the control group received conventional perioperative health education, while those in the intervention group received health education through an AI-based intelligent system. The primary outcome measure was the quality of recovery 24 hours after surgery, which was evaluated using the QoR-15 scale 24 hours after surgery; the secondary outcome measures included the time of first getting out of bed, the time of first passing gas, the incidence of postoperative diarrhea, the length of hospital stay, and the evaluation of system usability. Results: The health education system based on AI agents for patients with gallbladder stones during the perioperative period can improve the postoperative recovery quality of patients within 24 hours (P < 0.05), and the time for the first standing up and the first defecation of patients are both shortened (P < 0.05); however, there were no differences between the groups in terms of the incidence of postoperative diarrhea and the length of hospital stay (P > 0.05). Conclusions: The AI agent health education system provides patients with a health education tool that is readily accessible, provides continuous responses, and offers detailed guidance. It can serve as an auxiliary tool for nursing teams to provide health education and may facilitate patients’ postoperative recovery, providing a direction for the application of artificial intelligence technology in perioperative health education.

  • Background: Prospective payment systems (PPS) impose profound financial pressures on Intensive Care Units (ICUs), often fueling administrative anxiety that treating unavoidable high-acuity cases inevitably triggers catastrophic deficits. However, empirical evidence disentangling non-linear policy safety nets from actual structural inefficiencies remains scarce. Objective: To bridge this gap, this study leverages Explainable Artificial Intelligence (XAI) to decode the complex financial risk matrix under China's Diagnosis-Intervention Packet (DIP) system. Methods: We conducted a retrospective analysis of 5,258 ICU episodes at a major tertiary teaching hospital in China, utilizing deterministically linked clinical and financial datasets. Four algorithms were benchmarked to predict continuous financial deficits. The optimal Gradient Boosting Decision Tree (GBDT) model was integrated with SHapley Additive exPlanations (SHAP) to map localized marginal impacts and extract precise empirical zero-crossing thresholds. A prespecified 'pure ICU' sub-cohort sensitivity analysis was conducted to validate structural robustness. Results: GBDT significantly outperformed traditional linear models (R2 = 0.7674). SHAP dependence analysis revealed a sophisticated policy dichotomy. A profound "double zero-crossing" in the Cost Multiplier quantified a built-in safety net, providing a maximal protective trough at 2.75x. High-acuity metrics (RW > 4.93, LOS > 20.8 days) triggered similar financial buffering. Conversely, surpassing a rigid zero-crossing threshold of 23.51% in the Diagnostic & Lab Test Ratio induced a monotonic deficit surge (the "monitoring penalty"). Crucially, the model empirically debunked traditional cost-containment dogma: medical consumables played a surprisingly marginal role, whereas strategic interventions to optimize Length of Stay (LOS) averted systemic losses. These extracted thresholds demonstrated striking stability in the sensitivity overlay. Conclusions: The DIP framework operates not as a simplistic linear penalty for high-acuity care, but as a highly engineered matrix that buffers extreme physiological crises while strictly penalizing structural inefficiencies. By debunking the traditional consumables dogma, our XAI-derived blueprint empowers hospital administrators to transition from blind cost-cutting toward strategic clinical pathway optimization. Policymakers must introduce ICU-specific risk-adjustment coefficients to mitigate unintended monitoring penalties and safeguard value-based critical care.

  • Background: Drug-related seizures require timely recognition in complex treatment settings, but spontaneous-report associations, whole-molecule features, and network-structural clues are rarely examined together. Objective: To characterize drug-related seizures by integrating pharmacovigilance, high-dimensional chemical space, machine-learning applicability, and independent network and structural analyses. Methods: We curated FAERS data from 2004Q1 through 2026Q1 and defined seizures using SEIZURE plus historical CONVULSION. After deduplication and ingredient mapping, ROR, PRR, BCPNN, and MGPS were evaluated across drug-role and indication-exclusion specifications. We analyzed ring topology, Morgan fingerprints, and USRCAT space, then audited a benchmark XGBoost model using an applicability domain and refits. Separately, 12 neuro-oncology drugs entered network analysis; FAERS-positive drugs related to at least 2 hubs underwent docking and one 100-ns molecular-dynamics trajectory per complex. Results: The cohort contained 167,223 unique seizure-related reports. Auditing retained 68 of 844 eligible ingredients across specifications. Morgan and USRCAT spaces showed same-label enrichment in all 8 tests (empirical P=1/10,001; all Benjamini-Hochberg significant). XGBoost held-out accuracy was 0.783 and Active recall was 0.364. Eleven of 12 primary drug structures were in-domain, although 7 had unseen Murcko scaffolds. The 486 drug-associated targets and 3951 epilepsy-associated genes shared 236 genes, yielding 15 exploratory hubs. Independently selected everolimus and vincristine were screened against human HSP90AA1 and the 5EWM cross-species interface comprising Xenopus laevis GRIN1 chain C and human GRIN2B chain D. Four trajectories showed differential behavior; vincristine-HSP90AA1 had the most consistently low RMSD profile. Conclusions: Integrating reporting robustness, whole-molecule organization, model-scope auditing, and independent network-structural analysis produced a traceable strategy for prioritizing drug-context combinations and testable molecular hypotheses.

  • How rural Gypsy/Traveller women approach digital healthcare services as caregivers and care receivers: an interview study

    Background: Gypsy/Travellers are an ethnic group in the UK with poor health outcomes and multi-level challenges in services access, including potential digital barriers. In addition to these access and technical factors, gender also plays an important role in community healthcare in Gypsy/Traveller culture. This study explores the unique experiences of Gypsy/Traveller women in rural Scotland regarding how they engage with digital health services. Objective: To understand how Gypsy/Traveller women’s gendered roles and the use of technology shape each other in healthcare, and the implications for future services to be inclusive and supportive for this population. Methods: We recruited Gypsy/Travellers from a community organisation; participants were interviewed in small groups during community events. We drew on the concepts of intersectionality and ‘doing gender’ to explore Gypsy/Traveller women’s roles in healthcare and how technology impacted these roles. Results: Twenty-one participants (aged 16-71 years) were interviewed. Gypsy/Traveller women often manage healthcare for their family members while being healthcare recipients themselves; they have to navigate a hostile external environment due to systematic exclusion. Women are the main users of digital healthcare services, reflecting both their gendered roles and personal needs. Technology can help women juggle domestic responsibilities and provide space for them to discuss personal health, while at the same time, it may hinder relationship building with healthcare professionals and replace therapeutic in-person interactions. Conclusions: Out findings highlight how technology use in health is gendered when women are both caregivers and healthcare recipients. Technology holds the potential to support women but may also reinforce gender dynamics. To move towards a more equitable system, ethnic minority women should be included in service design; digital healthcare services meet the needs of women. The healthcare system needs to build up capacity to tackle discrimination and systematic exclusion of marginalised populations.

  • Mapping the Cancer Support Ecosystem on Reddit: A Computational Analysis of Community Connections, Supportive Care Needs, and Emotional Expression

    Background: Online communities provide large-scale, real-time insights into essential concerns and needs for people affected by cancer and have become an informal extension of the broader supportive care environment. Objective: This study aimed to identify the structure, content, and emotions in cancer-related Reddit discussions using large-scale data and computational methods. Methods: We analyzed 3,206,108 cancer-related posts published across 40 subreddits from 2010 to 2025. Shared authorship and cross subreddits movement networks were constructed to examine community connections. Guided by supportive care needs frameworks and validated clinical surveys, a Large Language Model classifies posts into five key needs: informational, physical, psychological, and practical. BERTopic was used to identify topics, and a RoBERTa-based model identified seven emotions: anger, disgust, fear, joy, neutral, sadness, and surprise. Temporal trends in emotion were examined, and differences across supportive care needs were assessed using chi-square tests and Cramér’s V. Results: Discussion volume steadily increased during the study period (2010-2025) and author activity distribution was highly skewed. The structure of the network characterized that r/cancer served as a central hub linking general, disease-specific, and support-oriented communities. The most prominent needs were informational (n=1,467,967, 45.79%), followed by psychological (n=584,306, 18.22%), physical (n= 508,149, 15.85%), social (n=331,953, 10.35%), and practical needs (n=104,722, 3.27%). Emotional expression changed over time, with sadness, anger, disgust, and joy decreasing, while fear and surprise increased. Emotion patterns varied across need discussions. Although neutral emotion was common overall, informational needs were characterized by fear and surprise; physical needs by disgust and sadness; psychological and social needs by joy and anger; and practical needs by sadness and anger. Conclusions: This study suggests that online cancer discussions capture both the content of unmet needs and the emotional context in which they are experienced. These results may inform patient-centered health communication, digital supportive care interventions, and future research using social media data to monitor cancer-related needs across the cancer continuum.

  • Evaluating Human–Large Language Model Interaction in Clinical Prediction Rule Calculation: Development and Validation Study

    Background: Clinical prediction rules (CPRs) are widely used in medicine, including for pediatric injury care. For retrospective validation, predictor variable extraction must be accurate and reproducible to avoid bias. However, extracting predictor variables from electronic health record (EHR) clinical notes is labor–intensive, difficult to scale, and prone to reviewer error. Objective: We aimed to evaluate a large language model (LLM) for its ability to calculate risk categories and extract predictor variables from emergency department (ED) physician notes in the EHR for four pediatric trauma CPRs. We also sought to characterize human–LLM interaction in clinical data extraction. Methods: We conducted a cross–sectional study of 750 de–identified pediatric emergency department visits in order to assess four published trauma CPRs in children: traumatic brain injury (TBI) in children younger than 2 years (TBI in younger), TBI in children older than 2 years (TBI in older), cervical spine injury (CSI), and intra–abdominal injury (IAI). The LLM performed zero–shot extraction using prespecified prompts (development cohort; n=300 visits), and performance was compared with expert reviewer extraction methods (validation cohort; n=450 visits). The primary outcome was LLM performance for CPR calculation and predictor variable extraction, including non–inferiority to expert reviewers using a prespecified 7.5% margin. Secondary outcomes examined the role of human–LLM interaction in reference standard creation. These included the number of expert reviewer extraction errors identified by the LLM, the resulting changes in children's risk categories, and the percentage decrease in extractions requiring full physician review. Results: The LLM’s risk category calculation accuracy was 91% (95% CI 85-96) for TBI in younger children, 97% (95% CI 94-100) for TBI in older, 95% (95% CI 91-98) for CSI, and 98% (95% CI 97-100) for IAI. The LLM’s risk category calculation accuracy and sensitivity were non–inferior to expert reviewers for TBI in older children, CSI, and IAI. LLM data extraction accuracy and sensitivity were also non–inferior to expert reviewers for 25 of 27 (93%) predictor variables. During reference standard creation, the LLM identified 34 human extraction errors (20% of corrections made), updated risk categories for 11 children (2.4%), and reduced physician review burden by 72%. Conclusions: An interactive human–LLM pipeline achieved expert–level performance for pediatric trauma CPR calculation and predictor variable extraction. Iterative prompt development improved reference standard quality and reduced physician review burden. These results also support a framework for human–LLM interaction spanning autonomous extraction, supervised screening, error identification, and consensus building. Our findings suggest that clinical data extraction is best viewed as an interactive process in which experts and LLMs contribute complementary strengths across multiple stages of extraction.

  • Background: Mental health-related content has been increasingly visible on Chinese social media among young people, including the expression of negative emotions and personal experiences of mental health issues, treatment and recovery, etc. Such content could reach a wide audience and influence their understanding and responses to mental health issues. However, previous studies have mainly focused on highly sensitive topics such as self-harm and suicide with occasional posts included, or examined discrete stages of content creation, while limited evidence can be found treating content creation as a connected process that involves motivation, creation, and reaction to feedback from the perspective of young creators. Objective: This study aims to explore the whole process of repeated content creation covering broader mental health topics, figuring out why Chinese young creators post, how they experience it, and what consequences they perceive. Methods: This study purposively recruited content creators aged 15 to 34 years from major Chinese social media platforms who have posted at least 5 posts about emotional distress, psychiatric symptoms,mental health treatment, recovery, psychological counseling, self-harm or suicide within 12 months. Online semi-structured interviews were conducted from December 2024 to March 2026 and analyzed using reflexive thematic analysis. Results: A total of 20 women completed interviews lasting 39 to 64 minutes, with a mean (SD) age of 21.8 (3.6) years. Fourteen participants (70.0%) primarily used Rednote. All participants reported at least one mental disorder diagnosis, including depressive disorder (n=10, 50.0%), bipolar disorder (n=7, 35.0%) and an anxiety and depressive state (n=3, 15.0%). Three themes were identified: 1) participants posted to meet emotional needs and made mental illness experience meaningful; 2) audience feedback shaped emotional experiences and subsequent creation and 3) disclosing mental illness required emotional effort and management of offline recognition, conflict, harassment, and unwanted contact. Conclusions: Repeated mental health-related content creation can provide connection and sense of value for young creators while also exposing them to unexpected visibility, conflicts and emotional burden through audience feedback. Social media platforms should strengthen appropriate content moderation with greater control over visibility and unwanted interactions, while family members could provide nonjudgmental offline support by acknowledging creators’ online experiences and respecting their disclosure boundaries.

  • Background: Delayed Code Stroke activation contributes to worse outcomes in acute stroke. Emergency Department (ED) triage notes contain free-text clinical information that could enable automated, real-time pathway activation. We evaluated the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting Code Stroke criteria from ED triage notes. Objective: The aim of this study was to evaluate the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting institutional Code Stroke criteria from ED triage notes and compare a single prompt to chain prompt performance. Methods: A retrospective cross-sectional study was conducted at Monash Medical Centre, Melbourne, Australia. De-identified triage notes from 3,023 ED presentations over a one-month period (September–October 2023) were analyzed. The pipeline applied sequential passes for translation, stroke symptom identification, mimic exclusion, baseline functional status, temporal window classification, and symptom resolution. Six locally deployed language models were evaluated. Performance was assessed against two reference standards: neurologist-labeled diagnosis and documented ED Code Stroke activation. Primary outcomes were sensitivity and specificity; secondary outcomes included PPV, NPV, and Gwet's AC1. Reliability of the neurologist reference standard was assessed by blinded independent re-review of a stratified random sample of 200 presentations by a second neurologist. Results: Of 3,023 presentations, 136 were neurologist-labeled positive. Agreement between the primary and a blinded second neurologist on a 200-note reliability sub-sample was almost perfect (raw agreement 95.0%, Cohen's κ 0.900, 95% CI 0.838–0.959). The cohort included 140 ED Code Stroke activations (median age 69, IQR 56–81 years), of whom 83 (59.2%) had confirmed stroke diagnosis. Sixteen patients (11.4%) underwent endovascular clot retrieval and 4 (2.9%) received thrombolysis. The best-performing model (Qwen 2.5 14B) achieved sensitivity 0.890 (95% CI 0.826–0.932), specificity 0.993 (0.989–0.996), PPV 0.858 (0.791–0.906), and NPV 0.995 (0.991–0.997). Pairwise McNemar testing demonstrated statistically superior overall accuracy for Qwen 2.5 14B over Llama 3.1 8B, Phi-4 14B, and Mistral 3 14B (all p<0.001 after Holm correction), with no significant difference detected versus Nemotron-Nano-12B-v2 or Qwen 3 14B. Conclusions: A locally deployed language model demonstrates acceptable sensitivity and specificity for automated Code Stroke identification from free-text triage notes. Performance was comparable across the two best models, suggesting that capable open-weight models in this parameter range may be sufficient to proceed with ongoing internal testing and external validation. The pipeline operates without internet connectivity or model retraining on patient data, supporting feasibility for real-world ED integration. Clinical Trial: ACTRN12626000596303

  • Quality of Information About Attention-Deficit/Hyperactivity Disorder on Social Media: a Scoping Review

    Background: There are 5.66 billion social media user identities worldwide, and content about ADHD is popular, with 6.1 million and 5 million videos under the hashtag “#ADHD” on Instagram and TikTok respectively. Given misinformation is common on social media, it is important to establish the quality and accuracy of information about ADHD on social media. A scoping review was chosen to explore the scope and nature of research on this topic, and to facilitate comparison across different social media sites and different research methods. Objective: This review examined the methods used by researchers to assess the quality of ADHD-related information on social media. Specifically, it aimed to identify the methodologies researchers used to assess the quality of information, synthesise the existing evidence, and identify areas and issues requiring more research. Methods: The scoping review was conducted in accordance with the PRISMA extension for scoping reviews (PRISMA-ScR). Three databases were searched: Medline, Embase, and PsycInfo. The search was limited to English language and papers published from 2009 onwards. The search strategy included the names of major social media platforms. Studies were eligible if they evaluated the quality of ADHD-related information available on social media. Ten studies met the inclusion criteria. Of these, 9 were conducted in high-income countries. Data was extracted independently by 2 reviewers who subsequently collaborated to synthesize the findings and identify the main themes and conclusions. Results: Four studies analysed TikTok, 1 study analysed both TikTok and Instagram, 4 studies analysed YouTube, and 1 study analysed 5 internet forums. Number of videos/posts analysed ranged from 45 to 159. Considerable variation was observed in the methods used to assess information quality. Four studies compared content to either the diagnostic criteria given in the DSM or a derivative of the DSM (ASRS-v1.1). Three studies used different generalised quality assessment tools designed for assessing health information. Four studies categorised content as “useful” or “misleading” using study-specific criteria. Despite methodological differences, most studies concluded that ADHD-related information on social media was of poor quality. Conclusions: The heterogeneity of methods used to assess quality limits the ability to draw broad conclusions about the quality of ADHD-related information available. Future research should adopt clearly defined quality assessment tools to improve consistency and comparability. Specifically, all reviewed studies had a limited sample size and 90% were conducted in high income countries. Moreover, the role of social media algorithms in presenting different content to different groups of people was not explored. Future research should include larger, more diverse samples of social media content and investigate how algorithm-driven content and geographical context affect the quality of information users encounter.

  • Advice and information social media users are seeking about prenatal cannabis use: Qualitative study of an anonymous online social forum

    Background: Self-reported use of cannabis amongst pregnant women rose from 1.5% to 5.4% from 2002-2020, despite strict guidelines against prenatal cannabis use. Furthermore, the American College of Obstetricians and Gynecologists have established guidelines against prenatal cannabis use due to its adverse effects on infant and child development. Due to the illicit nature of the drug, patients may fear legal consequences and withhold discussion with their provider. Anonymous online communities have become a sanctuary for discussion about the topic. Objective: The objective of this study is to examine anonymous online narratives to find reasons and mechanisms of use of cannabis by pregnant women, through the website Reddit.com. In doing this, the study aims to provides perspective on patient choices notwithstanding provider recommendations. Methods: Perinatal cannabis use discussion data was extracted from Reddit.com between March 2023 and June 2025 utilizing the software service, quid.com. Of these messages, data from the subreddit, r/pregnant, was extracted, resulting in a sample of 530 posts. Using the qualitative data analysis software, ATLAS.ti, an iterative content analysis was conducted. Results: Two major themes emerged from this analysis: (1) reasons for use and (2) modes of use. Reasons for use were coded under lack of perceived harm in child development (n=104), symptom relief (n=66), ineffectiveness of prescription medications (n=13), and use for previous diagnoses (n=21). Modes of use include a variety of cannabis formulations, such as THC (n=66), and CBD (n=23), through various means, such as smoking (n=189) and vaping (n=18), and different frequencies of use, from regular use (n=17) to an isolated event (n=1). Conclusions: Online discussions reveal that individuals frequently framed prenatal cannabis use as palliative treatment and harm reduction. Patterns of use ranged from use of unspecified cannabis formulations to specific compositions and mostly consumed through inhalation. Findings highlight gaps between clinical guidelines and patient perceptions, underscoring the need for non-judgmental, evidence-informed clinical counseling.

  • Artificial Intelligence for the Non-Invasive Diagnosis of Pulmonary Hypertension: A Systematic Review and Network Meta-Analysis

    Background: Pulmonary hypertension (PH) requires invasive hemodynamic confirmation, but right heart catheterization (RHC) is not always accessible. Noninvasive imaging is central to triage, yet the comparative diagnostic performance of conventional and artificial intelligence (AI)-enhanced modalities remains uncertain. Objective: To compare the diagnostic accuracy of conventional and AI-enhanced noninvasive imaging modalities for PH using RHC as the reference standard and to determine whether AI-enhanced approaches provide complementary value for clinical triage. Methods: In this systematic review and Bayesian diagnostic test accuracy network meta-analysis, PubMed, Embase, Web of Science Core Collection, and the Cochrane Library were searched from inception to March 29, 2026. Eligible studies enrolled adults with suspected or confirmed PH, evaluated imaging or imaging-derived AI models, used RHC as the main reference standard, and provided sufficient patient-level data to reconstruct 2×2 diagnostic tables. Eight prespecified nodes were compared: conventional and AI-enhanced echocardiography, computed tomography/computed tomography pulmonary angiography (CT/CTPA), conventional and AI-enhanced cardiovascular magnetic resonance (CMR), chest radiography-AI, and multimodal AI fusion. A Bayesian diagnostic test accuracy network meta-analysis estimated sensitivity, specificity, and diagnostic odds ratios, with ranking, inconsistency, meta-regression, publication-bias, and sensitivity analyses. Risk of bias was assessed with PROBAST+AI, and certainty of evidence with GRADE. The protocol was registered in PROSPERO (CRD420261431299). Results: Thirteen retrospective studies contributed 27 validation datasets; 22 used internal validation and 5 used external validation. CMR-AI had the highest pooled sensitivity (0.92; 95% credible interval [CrI] 0.84-0.97) but lower specificity (0.59; 95% CrI 0.31-0.82). Conventional echocardiography had the highest pooled specificity (0.92; 95% CrI 0.84-0.97) and diagnostic odds ratio (30.08; 95% CrI 10.77-66.29). CT/CTPA-AI showed higher sensitivity than conventional echocardiography (absolute difference 0.14; 95% CrI 0.00-0.29) and a higher relative diagnostic odds ratio than conventional CT/CTPA (2.21; 95% CrI 1.03-4.82). No modality simultaneously maximized sensitivity and specificity. Sensitivity analyses generally preserved the main trade-off pattern, although estimates for sparse nodes, particularly chest radiography-AI, were unstable. Meta-regression did not identify significant moderators. Evidence certainty ranged from high to very low, with very low certainty for multimodal AI fusion and concerns related to retrospective designs and frequent internal validation. Positive and negative posttest probabilities varied across nodes and assumed prevalence scenarios from 5% to 50%. Conclusions: AI-enhanced imaging may have a complementary role rather than replace expert interpretation or RHC. CMR-AI may support sensitive rule-out assessment, conventional echocardiography may support rule-in assessment, and CT/CTPA-AI may provide opportunistic decision support when cross-sectional imaging is already available. These findings are hypothesis-generating; prospective, multicenter, externally validated studies with standardized thresholds, calibration, and clinical-impact evaluation are needed before the results can guide routine diagnostic pathways.

  • Promoting mental health using a mobile health application in adolescents with a chronic somatic illness: A randomised waitlist controlled trial

    Background: Over 17% of children in the Netherlands live with a chronic illness. They are at risk of developing mental health problems, affecting quality of life and health outcomes. Objective: This study evaluates a preventive, transdiagnostic mobile health intervention, the Grow It! app, aiming to enhance overall well being by creating emotional insight and reflection, whilst encouraging the use of adaptive coping strategies. Methods: Using a randomised waitlist controlled design, adolescents aged 10–18 years with a chronic somatic illness were recruited through a university medical hospital and randomised to a waitlist control group or an intervention group that used the Grow It! app. Follow up assessments were conducted immediately after the four week intervention period and three months later. After the final follow-up, the waitlist control group also used the Grow It! app and completed questionnaires immediately after and three months later. Primary outcome measures were symptoms of anxiety and depression. Data were analysed with linear mixed models on an intention-to-treat basis. Results: Between August 2021 and September 2023, 208 participants were enrolled (Mage=13.96, SD=2.20, 57.69% female), with 105 randomised to the Grow It! intervention group and 103 to the waitlist control group. At the overall group level, symptoms of anxiety (χ2(3, 497.01)=2.08, p=.1021) and depression (χ²(3, 497.01)=0.71, p=.5459) remained stable over time, indicating no significant change across the full sample. However, among adolescents who used the Grow It! app, adolescents with elevated anxiety/depressive at baseline demonstrated the greatest benefit, showing larger symptom reductions than those who initially reported low symptom levels. Conclusions: These findings suggest that the Grow It! app may offer meaningful benefits for adolescents with a chronic somatic illness experiencing elevated mood and anxiety symptoms, highlighting its potential value as a preventive and supportive tool in paediatric healthcare. Clinical Trial: International Standard Randomized Controlled Trial Number (ISRCTN) 17883961; https://doi.org/10.1186/ISRCTN17883961

  • RandIMI – Building a Randomization Service for Multicenter Studies: Development and Usability Study

    Background: Randomized controlled trials (RCTs) are the gold standard for evaluating medical interventions because randomization reduces bias and provides high-quality evidence. To further minimize confounding, participants are often stratified by prognostic variables before random assignment to treatment groups. As clinical trials are increasingly managed digitally, integrating randomization directly into hospital information systems and electronic data capture platforms can simplify trial workflows by eliminating the need for separate software or manual randomization lists. Furthermore, the growing prevalence of multicenter studies highlights the need for interoperable and flexible randomization services. Objective: Our objective was to provide a web service that handles the randomized assignments in clinical trials and integrates seamlessly into the existing digital infrastructure. By considering versatility in trail designs, we intended to ensure the capability to apply our service in various real-world clinical research trials. Methods: Different randomization algorithms are leveraged to ensure the quality of randomized assignments, such as restricted, blocked, and dynamic randomization. Evaluation was done by surveying clinicians who use RandIMI and its REDCap integration in their real-world clinical trials using the standardized System Usability Scale. Results: We present RandIMI, an open-source web-based randomization service designed for both standalone use and seamless integration into existing clinical research platforms via a REST API. Existing integrations include REDCap and the hospital information system ORBIS. RandIMI supports a wide range of trial designs, including multicenter studies, an unlimited number of treatment groups with configurable allocation ratios, and stratification by study site and categorical variables. Its flexible study model enables modifications during recruitment, allowing adaptation to evolving trial requirements. Comprehensive user management and an audit trail ensure data security and traceability. RandIMI further provides configurable static and dynamic randomization algorithms to support robust allocation while minimizing selection bias. RandIMI achieved a SUS score of 77.3 from a total of 24 participants, indicating good usability. To date, the system has been successfully deployed in 15 real-world trials demonstrating its applicability across diverse study designs. RandIMI is freely available as open-source software on GitHub: https://github.com/imi-ms/RandIMI. Conclusions: This study shows that the presented randomization service RandIMI is capable of supporting research studies by providing a robust and user-friendly method to conduct the randomization of participants. Its API and integrations into REDCap and ORBIS offer a simple yet powerful extension to the trial infrastructure. Although audit trails, deterministic assignments, and a recruitment history are provided, RandIMI is not certified for usage in the development and approval of medical products.

  • Caregiver-Informed Safety and System Requirements for Real-Time Feedback During Pediatric Inhaler Use: Qualitative Concept-Elicitation Study

    Background: Correct use of a pressurized metered-dose inhaler (pMDI), usually with a valved holding chamber, depends on use-related conditions that caregivers may not observe directly, including inhalation pattern, actuation-inhalation timing, mask seal, and device orientation. Many digital inhaler systems emphasize actuation or adherence. Systems that infer technique-related conditions from proxy signals may add real-time feedback but also create false reassurance if a positive state is displayed when evidence is inadequate. Caregiver expectations for such feedback and its failure states remain poorly characterized before prototype development. Objective: This study aimed to characterize caregivers' experiences of pediatric inhaler administration and translate responses to a proposed sensor concept into traceable candidate safety, system, and validation requirements. Methods: We conducted 23 semistructured concept-elicitation sessions by Zoom videoconference (19 individual and 4 dyadic sessions within households) with 27 caregivers of children who currently or previously used an inhaler. A 6-phase guide examined current practices, difficulties, prior digital-tool use, and responses to a screen-shared clip-on concept incorporating airflow-, timing-, seal-, and orientation-related sensing; on-device feedback; an optional child-facing app; and a caregiver dashboard. No working prototype was evaluated. Approximately 86,000 transcript words underwent hybrid directed and conventional qualitative content analysis using an 18-category codebook. A second researcher independently coded the complete 27×18 matrix (486 paired decisions) across 2 rounds. Consensus categories were translated into candidate requirements and assigned an earliest verification or evaluation locus. Counts are descriptive caregiver-voice tallies, not prevalence estimates. Results: Persistent uncertainty about whether administration conditions were adequate was the dominant lived experience (19/27), and existing passive cues did not resolve it (5/27). A simple, real-time on-device signal was widely endorsed (25/27), but willingness to consider the concept was conditional on avoiding false-positive reassurance (a “false green”; 9/27), low-friction operation during acute symptoms (12/27), minimal added bulk and technique change, and one or more of clinical validation, clinician recommendation, or a trial pathway (composite category; 23/27). Stances on gamification (20/27), dashboard cadence (18/27), aesthetics (24/27), and privacy (18/27) supported an optional, configurable software layer. Intercoder agreement was 94.0% (Cohen kappa=0.88). The analytic synthesis yielded a conceptual architecture separating observable proxy signals, an inference and data-quality gate, and positive, corrective, or uncertain feedback states. Conclusions: Caregivers prioritized immediate, interpretable feedback about validated use-related conditions, not proof that medication reached the lungs. The findings support conservative state assignment, an explicit uncertain state for indeterminate signals, and app-independent acute-use feedback. These are candidate requirements rather than validated device specifications; subsequent work should proceed through bench verification, simulated-use human factors evaluation, and prospective clinical or field validation.

  • Understanding lichen sclerosus through online narratives: A qualitative analysis

    Background: Lichen sclerosus (LS) is a chronic inflammatory dermatosis associated with significant physical and psychosocial morbidity. Despite established management guidelines, patients frequently experience diagnostic delays, inconsistent counseling, and uncertainty regarding disease progression. Online patient communities provide a unique opportunity to examine lived experiences that may not be fully captured in clinical settings. Objective: To characterize patient experiences, psychosocial concerns, and emotional burden related to LS through analysis of discussions within an online patient support community. Methods: We conducted a cross-sectional computational analysis of publicly available discussions from Reddit’s r/lichensclerosus community between 2017 and 2025. Posts and comments were analyzed separately using topic modeling, psychosocial multi-label classification, and emotion and affect analysis. Temporal trends in discussion topics and psychosocial themes were also evaluated. Results: A total of 1,313 posts and 659 comments were included. Discussion centered around illness uncertainty and fear of disease progression, accounting for 64.1% of posts and exhibiting the highest levels of emotional distress. Treatment-related concerns, particularly topical corticosteroid use, represented the second most common theme (21.2%). Psychosocial burden was substantial, with support-seeking behaviors identified in 86.97% of posts, alongside frequent discussion of sexual dysfunction and body image disturbance. Clinician mistrust demonstrated the strongest association with psychosocial burden across both posts and comments. Posts were characterized primarily by first-person narratives of distress and uncertainty, whereas comments focused on reassurance, shared experiences, and practical management strategies. Emotional burden was consistently higher in posts than comments, suggesting a potential buffering effect of peer support and community engagement. Temporal analyses demonstrated minimal changes in discussion patterns over the study period. Conclusions: Online communities serve as important spaces for validation, knowledge sharing, and informal care navigation among individuals with LS. These findings suggest that clinic-based perspectives may not fully capture the daily lived experiences and psychosocial challenges associated with the disease. Improved patient-centered education, clinician-patient communication, and integration of psychosocial support into LS care may better address patient needs. Leveraging real-world patient discourse offers valuable insight into patient priorities and may help inform more responsive and empathetic models of care. Clinical Trial: Not applicable.

  • Background: Studies comparing large language models (LLMs) with human qualitative analysis of English and Japanese clinical data report stronger performance on descriptive themes than interpretive themes. Evidence from Chinese nursing data remains limited to 1 or 2 models. Few studies have compared multiple heterogeneous models on one dataset while clearly separating analytic tasks. Objective: This study compared 6 global and China-developed LLMs with a human consensus reference using Chinese parent interviews about pediatric day surgery. We evaluated inductive theme discovery and blinded predefined code application, including repeated runs and exploratory cultural cases. Methods: We analyzed semistructured interviews with 21 parents at a Chinese tertiary hospital, comprising approximately 240,000 Chinese characters. In phase 1, 2 researchers independently applied the Colaizzi method, with adjudication by a third researcher. The resulting human consensus reference comprised 4 themes and 13 codes. The models were Claude Opus 4.8, ChatGPT 5.5, Gemini 3.1 Pro, Qwen 3.7 MAX, Zhipu GLM 5.2, and Doubao 2.0 Pro. They generated themes under identical prompts, and their outputs were compared qualitatively. In phase 2, each model applied the locked codebook to 100 meaning units from the first 5 transcripts. Models assigned blinded binary labels without iterative feedback in 3 independent same-day runs. The primary analysis used 3-run majority-vote labels. We calculated Cohen kappa, sensitivity, specificity, and F1 scores with participant-clustered bootstrap 95% CIs. We also report single-run kappa ranges and unit-level stability. Results: Ensemble agreement varied by theme and model. For perioperative fear management (56/100 positive units), 4 models had kappa values of 0.13-0.19. Claude and Qwen reached 0.67 and 0.69, respectively. Discharge readiness (5/100 positive units) yielded kappa values of 0.56-0.90 with wide CIs. Care-process optimization yielded values of 0.50-0.72. The median theme gradient persisted, but poor performance on implicit themes was not uniform across models. Three-run unit-level agreement ranged from 45% to 100%. ChatGPT produced same-day kappa values of 0.66, 0.13, and 0.10 for fear management. One model produced a degenerate over-labeling run, and only Claude remained consistently high. Four models showed high specificity but low sensitivity (0.14-0.21) for fear management. All 6 models missed 3 of the 13 human-derived codes. Three exploratory cultural cases suggested that China-developed models more often reflected local context, although these findings are not generalizable. Conclusions: LLM performance varied jointly by task type, theme explicitness, and model. Poor performance on implicit themes should not be treated as a uniform property of all LLMs. Performance differed substantially across models, and only 1 model maintained consistently higher agreement across repeated runs. Future evaluations should include repeated runs and report between-run variability. Majority voting may provide one robustness strategy, but human arbitration remains essential. The bidirectional unique contributions observed here are a precondition for complementary team performance, not evidence of such performance. We therefore propose prospective evaluation of a human-supervised workflow against human-only and model-only comparators.

  • Robotic versus Laparoscopic Surgery for Colonic Diseases: An Overview of Systematic Reviews and Meta-Analyses

    Background: Colonic diseases frequently require minimally invasive procedures. Laparoscopy has become the standard minimally invasive approach but is limited by 2D vision, magnified hand tremors and poor ergonomics. Robotic surgery was introduced to overcome these shortcomings, yet it suffers from longer operation time and higher costs. Multiple meta-analyses comparing the two techniques have yielded inconsistent perioperative and oncological findings, with narrow disease coverage and limited generalizability. This overview synthesizes current evidence to inform clinical decision-making. Objective: The aim of this overview of systematic reviews was to systematically evaluate the advantages and perioperative outcomes of robotic surgery compared with laparoscopic surgery for colonic diseases, and to provide evidence-based guidance for clinical decision-making. Methods: PubMed, Embase, and Web of Science were searched from inception to January 1, 2026 for systematic reviews and meta-analyses comparing robotic and laparoscopic colonic surgery. After screening against predefined inclusion and exclusion criteria, the PRISMA 2020, AMSTAR 2.0, and GRADE tools were used to assess the reporting quality, methodological quality, and level of evidence of the included reviews. Outcomes including conversion to open surgery, operative time, blood loss, complications, length of hospital stay, oncological outcomes, medical costs, and mortality were analyzed.This overview was registered on PROSPERO (CRD420261415344). Results: Twenty-four English-language meta-analyses were included. Quality assessment showed that most reviews had some reporting deficiencies; none provided a list of excluded studies or funding information. The overall evidence was predominantly low to very low. Outcome analysis indicated that robotic surgery was associated with lower conversion rate to open surgery, less intraoperative blood loss, shorter postoperative hospital stay, and shorter time to first flatus. However, operative time was significantly longer for robotic surgery. No significant differences were found between the two groups for anastomotic leakage, incisional infection, postoperative ileus, number of lymph nodes harvested, or all-cause mortality in most studies. Robotic surgery incurred higher medical costs. Conclusions: Robotic colonic surgery offers superior minimally invasive benefits and better postoperative recovery, making it suitable for patients with technically demanding conditions. Laparoscopic surgery has shorter operative time and lower costs, making it more appropriate for routine colonic lesions and resource-limited settings. The overall quality of the available evidence is limited, and more high-quality, large-sample randomized controlled trials are needed.

  • Redesigning the Optical Viewing Environment for Children’s Digital Learning With a Foldable Distant-Image Display: Randomized Controlled Trial

    Background: As digital and AI-enabled learning become embedded in children’s education, the relevant health question is not only how much screen time children accumulate, but whether the optical viewing environment can be redesigned. Conventional tablets are generally viewed at short distances and require sustained accommodation and convergence. Earlier distant-image systems were predominantly fixed desktop devices, limiting storage, portability, and deployment in homes and schools. A foldable display that preserves distant-image viewing could alter the optical exposure associated with digital learning without changing course content. Objective: To compare short-term ocular responses after an identical 40-minute digital learning task delivered through an S1 foldable distant-image display or a conventional tablet in children. Methods: This single-center, prospective, parallel-group randomized controlled trial enrolled 48 children aged 6 to 16 years and allocated them in a 1:2 ratio to a conventional tablet group (n=16) or an S1 group (n=32). Both groups viewed the same online course and video content for 40 minutes. The tablet was viewed at approximately 40-50 cm; the S1 presented the visual content as a virtual image at approximately 6 m. The primary outcome was the change in noncycloplegic spherical equivalent (SE). Key secondary outcomes were changes in subfoveal choroidal thickness and a 1-to-5 visual fatigue composite score adapted from 16 Computer Vision Syndrome Questionnaire symptom domains. Exploratory outcomes were logMAR visual acuity and short-term axial length. Baseline-adjusted linear models included age and used HC3 heteroskedasticity-consistent standard errors. Results: The mean ages were 11.1 (SD 2.6) years in the tablet group and 10.2 (SD 2.2) years in the S1 group. Mean SE changed by -0.27 D after tablet viewing and +0.10 D after S1 viewing; the adjusted between-group difference was +0.38 D (95% CI 0.27-0.49; P<.001). Mean choroidal thickness decreased by 11.1 μm in the tablet group and increased by 11.2 μm in the S1 group; the adjusted difference was +22.0 μm (95% CI 14.8-29.2; P<.001). Visual fatigue increased by 0.69 points in the tablet group and decreased by 0.22 points in the S1 group; the adjusted difference was -0.62 points (95% CI -0.87 to -0.38; P<.001). The adjusted differences in logMAR visual acuity and axial length were -0.049 (95% CI -0.072 to -0.026; P<.001) and -0.012 mm (95% CI -0.023 to -0.001; P=.041), respectively. No device-related adverse events occurred. Conclusions: During a single 40-minute digital learning task, the foldable distant-image display produced more favorable short-term SE, choroidal, and visual-fatigue responses than a conventional near-viewed tablet. These findings provide randomized mechanistic evidence that desktop distant-image optics can be translated into a foldable display modality for viewing-centered digital learning. They do not establish long-term myopia prevention, control of axial elongation, or full functional equivalence to a tablet computer. Clinical Trial: Chinese Clinical Trial Registry, ChiCTR2100047059; registered June 7, 2021;

  • Inter-Agent Information Sharing in Multi-Agent Medical Diagnosis: A Controlled Offline Repeated-Measures Evaluation

    Background: In multi-agent medical diagnostic systems, diagnostic tasks are distributed across specialized agents and connected through inter-agent information sharing. Whether different inter-agent information-sharing conditions improve final diagnostic performance or primarily alter intermediate diagnostic outcomes remains unclear. Objective: To determine whether different inter-agent information-sharing conditions affect final diagnostic performance and intermediate diagnostic outcomes within a standardized multi-agent diagnostic workflow. Methods: We conducted an offline repeated-measures evaluation of 297 cases from three datasets using three large language models (LLMs). Each case–model pair was tested under 11 conditions: the single-agent baseline (SA); complete structured sharing (C1); structured sharing with explicit reasoning (C2); controlled field-level sharing (C3); minimal R3 coordination (C4); and six exploratory agent-role ablation conditions (A1–A6), yielding 9,801 case–model–condition evaluations. Prespecified analyses compared C1–C4 with SA and evaluated C2 versus C1, C3 versus C1, and C4 versus C3; C3 served as the reference for A1–A6. Clinical equivalence of the R2 candidate diagnoses, R5 final diagnosis, and final top-3 diagnoses was independently assessed by two physicians blinded to the base LLM and experimental condition, with disagreements adjudicated by a third physician. Accuracy was the principal outcome; top-3 hit rate, MRR@3, R2 differential diagnosis coverage, and token use were secondary outcomes; diagnostic error propagation and correction outcomes and agent-level output measures were exploratory. Results: C1 had the highest observed accuracy among C1–C4 (60.47% vs 59.36% for SA), but no multi-agent condition showed a statistically supported improvement in accuracy, top-3 hit rate, or MRR@3. Structured sharing with explicit reasoning (C2) increased mean tokens per case relative to C1. In exploratory analyses of agent-level output measures, controlled field-level sharing (C3) increased the R4 conflicting-evidence and coverage-gap count relative to C1 without statistically supported changes in diagnostic error propagation and correction outcomes or final diagnostic performance. Each agent-role ablation condition reduced mean tokens per case by 5,743–12,784 relative to C3 (all Holm-adjusted P=.02) without statistically supported differences in final diagnostic performance; removing R1 alone or with R3 also reduced the R4 count (both Holm-adjusted P=.01). Conclusions: Within the evaluated datasets, base LLMs, prompts, and standardized multi-agent diagnostic workflow, the inter-agent information-sharing conditions altered selected intermediate outcomes and token use but did not produce statistically supported gains in final diagnostic performance. Exploratory agent-role ablation findings do not establish noninferiority, equivalence, or general dispensability of any role. Further evaluation in interactive clinical settings is required before extending these findings beyond offline diagnostic tasks.

  • Beyond Sycophancy: Experimental Study of Large Language Model Methodology Review

    Background: Large language models (LLMs) have been demonstrated to exhibit sycophantic behavior, agreeing with users rather than corrected them, which may reinforce user belief over accuracy. Whether this compromises LLM appraisal of flawed research methods is unknown. Objective: This study evaluated whether LLMs endorse methodologically flawed clinical research excerpts as sound, whether they correctly pass sound excerpts, and whether stated user confidence or authority alters these behaviors. Sycophantic false reassurance was hypothesized to occur and to increase with user confidence. Methods: Twenty clinical research methods excerpts were developed: ten containing a single prespecified major methodological flaw and ten matched flaw-corrected twins. Excerpts were submitted to two commercial LLM vendors (OpenAI, Anthropic) as independent structured review requests. The primary outcome was false reassurance, defined as endorsing a flawed excerpt as sound. The secondary outcome was clean overflagging, defined as recommending major revision of a flaw-corrected excerpt. The application program interface (API) experiment crossed 20 excerpts by 2 vendors, 7 prompt framings, 3 stated expertise levels, and 5 replicates (4200 responses). A targeted 200-response substudy was collected through the consumer ChatGPT and Claude interfaces with expertise fixed at intermediate. Results: False reassurance was rare (3/2100 flawed responses, 0.14%; 0/1050 Anthropic, 3/1050 OpenAI, 0.29%) and did not occur in the commercial product substudy (0/40). Clean overflagging occurred in 14.5% (152/1050) of Anthropic and 48.8% (509/1044) of OpenAI clean responses (OR 5.70, 95% CI: 2.92 – 11.10). This difference persisted after sequential exclusion of the most frequently flagged excerpts. Stated expertise was not associated with either outcome. The overall prompt-framing effect was not significant (p = .26), though authority-pressure (OR 1.81, 95% CI: 1.10 – 2.97, p = .02) and high-user-confidence (OR 2.74, 95% CI: 1.05 – 7.14, p = .04) were associated with higher odds of clean overflagging than neutral framing. Clean overflagging was substantially higher through consumer products than the API for both vendors. Conclusions: Sycophantic false reassurance was essentially absent and asserting user confidence or authority did not induce it. The dominant behavior was instead recommending revision of methodologically sound excerpts, which varied markedly by vendor and by whether the model was accessed through a consumer product or an API. Given identical excerpts produced different verdicts across products and surfaces, model-generated revision recommends should prompt human evaluation of the stated concern rather than automatic acceptance that a flaw exists. Clinical Trial: NA

  • Background: Obesity is a critical global public health crisis with high prevalence in China, and glucagon-like peptide-1 receptor agonists (GLP-1 RAs) have emerged as novel weight-management agents. In 2024, China implemented official weight-management policies, making it essential to analyze public discourse about anti-obesity drugs on Chinese social media to support rational promotion and public health practice. Objective: This study aimed to analyze public discussions regarding emerging anti-obesity drugs on the Chinese social media platform Weibo, and to evaluate the changes in public discourse patterns following the implementation of China's 2024 weight-management policy. Methods: This cross-sectional study analyzed 25,733 valid Weibo posts about novel anti-obesity drugs from December 2017 to December 2025, using Latent Dirichlet allocation for topic modeling and interrupted time-series analysis to evaluate changes in public discussion trends before and after the policy. Results: We identified three major themes—Medication Usage Experience, Pharmaceutical Development Dynamics, and Medication Science Communication—and 11 subtopics. Most posts came from general users and focused on semaglutide. After the policy, the proportion of Medication Usage Experience discussions decreased from 11.76% to 7.10%, the proportion of Pharmaceutical Development Dynamics remained relatively stable (from 50.51% to 50.01%), while the proportion of Medication Science Communication increased from 37.73% to 42.89%. Interrupted time-series analysis further revealed that Medication Science Communication declined sharply (coefficient = −29.16; 95% CI: −41.90 to −16.42; P < .01) and then rose significantly (coefficient = 0.05; 95% CI: 0.04 to 0.07; P < .01). Conclusions: After policy intervention, discussion initially decreased but subsequently showed a relative increase in science-oriented content, suggesting a potential shift in public focus. These findings thus indicate public responses to obesity interventions and support scientific promotion of anti-obesity drugs and weight-management strategies.

  • Measuring Content Engagement and Associated Outcomes in a Digital Health Vaccine Intervention for Black Young Adults: A Paradata Analysis

    Background: Digital health interventions (DHIs) hold promise for improving health outcomes. Tough Talks COVID-19 (TT-C) was a culturally tailored, community-informed DHI designed to support Black young adults (YAs) in the US South in making informed COVID-19 vaccination decisions. Objective: To understand how user engagement, captured through paradata, was associated with improved vaccine attitudes in the TT-C DHI. Methods: We analyzed paradata from 299 Black YAs in Alabama, Georgia, and North Carolina who used the TT-C DHI in a randomized controlled trial. The DHI included 62 self-paced activities that were analyzed using rapid content analysis to identify the specific vaccine uptake-related domain they covered (hesitancy, confidence, conspiracy beliefs, and knowledge). Intervention engagement was defined as number of activities and time spent in the DHI. We used regression models to evaluate associations between intervention engagement and changes in vaccine-related measures after three months of use. Results: Overall engagement with the DHI was high.. Greater engagement with the DHI, both in total time and number of activities completed, was associated with larger improvements for all four vaccine outcomes at three months after download (P < 0.05). The effects of the TT-C DHI were amplified when participants engaged in activities that were directly related to the domain of interest, particularly with conspiracy beliefs, which had an effect size three times larger for domain-specific engagement (β: -0.067 [95% CI: -0.130, -0.005]) compared to domain-agnostic engagement (β: -0.019 [95% CI: -0.035, -0.002]) Conclusions: Our findings underscore the value of using paradata to assess DHI effectiveness and highlight that the type of content engaged with is associated with change. These insights can guide future DHI design and real-time adaptive interventions to maximize impact. Clinical Trial: ClinicalTrials.gov NCT05490329; https://clinicaltrials.gov/study/NCT05490329

  • Diagnostic Performance of Deep Learning for Orbital Fracture Detection on CT: A Systematic Review and Meta-Analysis

    Background: Orbital fractures are common facial injuries, yet subtle fractures are frequently missed on CT due to visual fatigue and subjective interpretation. Missed fractures can lead to preventable complications including enophthalmos, diplopia, and permanent visual impairment if surgical intervention is delayed. Deep learning offers a promising solution for automated detection, but evidence on its diagnostic accuracy remains fragmented and has not been systematically synthesized. Objective: We sought to evaluate the overall diagnostic performance of deep learning (DL) for detecting orbital fractures on computed tomography (CT), and to explore heterogeneity by comparing analysis units (slice-level vs. non-slice-level). Additionally, we aimed to conduct exploratory evaluations of neural network architectures (CNNs vs. Transformers) and fracture subtype classification (trap-door vs. depressed). Methods: A systematic literature search was conducted across PubMed, Embase, Web of Science, Cochrane Library, and Scopus up to July 6, 2026. Methodological quality was assessed using QUADAS-2 and PROBAST+AI. For the primary diagnostic performance and unit-of-analysis comparisons, pooled sensitivity, specificity, and the area under the curve (AUC) of the optimal independent models were calculated using a bivariate random-effects model, strictly ensuring no patient overlap. Conversely, to prevent optimism bias, we exhaustively extracted data from all available models to conduct an exploratory comparison between CNN and Transformer architectures, acknowledging the inherent patient overlap. Furthermore, due to the availability of only a single study for subtype classification (trap-door vs. depressed), diagnostic metrics were solely extracted without meta-analytic pooling. This review was prospectively registered in PROSPERO (CRD420261442640). Results: Five studies encompassing 31,546 diagnostic instances were included. Without patient overlap, the optimal pooled DL models demonstrated excellent overall diagnostic performance with moderate certainty of evidence, yielding a sensitivity of 0.96 (95% CI: 0.92-0.98), specificity of 0.96 (95% CI: 0.87-0.99), and an AUC of 0.98 (95% CI: 0.95-0.99). Subgroup analysis indicated that slice-level evaluations artificially inflated specificity compared to non-slice-level evaluations (0.96 vs. 0.83, P < 0.001). In the exhaustive exploratory analyses containing patient overlap, CNNs maintained a robust sensitivity of 0.88, whereas Transformers were compromised (0.41); additionally, algorithms showed high accuracy (0.83-0.92) in classifying trap-door versus depressed subtypes within the single available cohort. Conclusions: DL demonstrates excellent diagnostic accuracy for detecting orbital fractures on CT, showing great potential as a triage tool. However, slice-level assessments can artificially inflate specificity. Exploratory comparisons suggest CNNs may outperform Transformers and show potential in identifying critical trap-door subtypes, though these specific findings are limited by inherent patient overlap and lack of multicenter external validation. Future large-scale prospective validations adhering to standardized guidelines are urgently required to verify these findings.

  • Background: Breast cancer is among the most common malignancies worldwide. Despite its high public awareness and advances in diagnosis and treatment, many patients continue to experience substantial unmet informational needs. Patients and their relatives frequently seek additional information online, making infodemiological analyses a valuable tool for identifying public concerns, knowledge gaps, and potentially underserved patient populations through online search behavior. Objective: This study aimed to characterize breast cancer–related online search behavior in Germany using infodemiological methods. Methods: Breast cancer-related online search behavior in Germany was analyzed using Google Trends to assess monthly search trends from January 2016 to December 2024 and Google Ads Keyword Planner to identify and systematically analyze associated search terms. A weighted scoring system based on search frequency was applied to further investigate treatment-related side effects and rare patient subgroups. Results: The lay term showed substantially higher search interest than the medical term indicating a clear preference for non-medical terminology. High search interest was observed for treatment-related side effects, particularly alopecia and pain. Among special patient groups, male breast cancer accounted for the largest proportion of related search queries despite its low incidence. Conclusions: Future patient information strategies should place greater emphasis on lay terminology to better reflect real-world search behavior. The remarkably high search volume related to male breast cancer is particularly notable given the generally lower health information–seeking behavior among men and may indicate barriers to access medical support. Infodemiological analyses may help identify underserved groups and support more patient-centered communication strategies.

  • Shame on Who? #FastTailedGirl Discourse on Twitter and Its Implications for Black Girls' Sexual and Reproductive Health: A Qualitative Analysis

    Background: In the United States, Black girls and women are disproportionately affected by sexually transmitted infections (STIs), HIV, and sexual violence [1-3]. The "fast-tailed girl" (FTG) narrative, a gender-specific pejorative suggesting a girl is intentionally demonstrating sexual behaviors reserved for adults, has persisted in Black communities for generations and intersects with broader patterns of adultification bias that deny Black girls the protections of childhood [4-6]. Social media platforms have become significant sites where this harmful narrative is both perpetuated and contested [7,8]. Objective: This study examines Twitter discourse surrounding the #FastTailedGirl hashtag to understand how this narrative manifests online, its implications for Black women's and girls' sexual development and health, and how community members resist this harmful trope. Methods: Using Brandwatch, we retrieved 34,485 English-language tweets containing #fasttailedgirl and related permutations posted between 2011 and 2021, with retweets excluded. Tweets identified as originating outside the United States were excluded. Tweets with no country code available were retained because missing geographic metadata do not indicate that a tweet originated outside the United States. This yielded an eligible dataset of 33,099 tweets, from which a 20% random sample (N=6,899) was selected for qualitative analysis. We analyzed this sample using reflexive thematic analysis informed by intersectionality and reproductive justice frameworks [9,10]. After tweets coded as irrelevant or lacking sufficient context (code 99; n=678) were removed, the final thematic analytic sample included 6,221 tweets. Results: Two overarching themes emerged: (1) a culture of acceptance, encompassing blaming and shaming of Black girls, sexualization and predatory behavior, denial of childhood, generational transmission, and long-term consequences including under-reporting of sexual assault; and (2) resistance through provocation and inquiry, including myth-busting, activism and advocacy, and reclaiming self and body autonomy. Conclusions: The FTG narrative represents a significant public health concern that may reflect and reinforce conditions shaping sexual and reproductive health inequities among Black women and girls. Findings point to an urgent need for culturally responsive digital and community-based interventions that counter stereotype messaging, create safe spaces for empowerment, and support healing from trauma. Clinical Trial: Not Applicable

  • The role of Artificial Intelligence in interventions to reduce physical inactivity: A systematic review

    Background: Prevalence of physical inactivity is high and increasing across high-income countries, and interventions to increase physical activity can include behaviour change delivered digitally. Artificial intelligence (AI) is a broad field encompassing various techniques, which include algorithms that learn from data to perform automated tasks without explicit human programming, and could potentially increase the effectiveness of digital behaviour change interventions through personalisation and reducing barriers to engagement. Objective: This review aims to summarise the evidence for the effectiveness of AI-assisted interventions to reduce physical inactivity. Methods: We conducted a systematic review to identify and summarise evidence from randomised controlled trials (RCTs) of AI-assisted interventions for reducing physical inactivity. Eligible trials were RCTs that reported results of an AI-assisted public health intervention for reducing physical inactivity in a high-income country. We searched Medline (Ovid), Embase (Ovid), Web of Science (Core collection), and Scopus for relevant trials published between 2010 and 14 January 2025. We also searched for reviews of public health interventions for physical inactivity published between 2023 and 14 January 2025, and extracted all references from relevant reviews for screening. We conducted forward and backward citation searching on all included trials (dates of searches: September 2025 to June 2026). Screening for trials was conducted independently by two reviewers. Two reviewers independently assessed risk of bias using the Cochrane Risk of Bias 2 tool. As the included trials were heterogeneous in terms of interventions, outcomes, and timepoints, we synthesised the results narratively. Results: We included 17 trials (comprising 48 reports and 3,282 randomised participants). Five trials estimated the effectiveness of apps with chatbots for reducing physical inactivity (one with some concerns of bias, four with high risks of bias), with little evidence to suggest that chatbots increase physical activity. Twelve trials estimated the effectiveness of selection of motivational or other messages using recommender systems, reinforcement learning, machine learning, or case-based reasoning for reducing physical inactivity (four with some concerns of bias, eight with a high risk of bias), with little evidence to suggest an increase in physical activity generally, though some evidence to suggest an increase in step count specifically. All trials had relatively few participants, so results were generally imprecise. There were no trials using large language models. Conclusions: There is no strong evidence of a beneficial effect of AI-assistance in public health interventions for reducing physical inactivity in high-income countries, though AI-assisted interventions may increase step count. Future research should better describe public health interventions that use AI and embed equity considerations into their design and analysis to ensure already disadvantaged groups are not harmed further by the adoption of AI-assisted interventions in public health. Clinical Trial: PROSPERO CRD42025642339

  • The role of Artificial Intelligence in interventions to reduce alcohol consumption and smoking: A systematic review

    Background: Prevalence of alcohol consumption and smoking is high across high-income countries, and interventions to decrease either can include behaviour change delivered digitally. Artificial intelligence (AI) is a broad field encompassing various techniques, which include algorithms that learn from data to perform automated tasks without explicit human programming, and could potentially increase the effectiveness of digital behaviour change interventions through personalisation and reducing barriers to engagement. Objective: This systematic review aims to summarise the evidence for the effectiveness of AI-assisted interventions to reduce alcohol consumption and smoking. Methods: We conducted a systematic review to identify and summarise evidence from randomised controlled trials (RCTs) of AI-assisted interventions for reducing alcohol consumption or smoking. Eligible trials were RCTs that reported results of an AI-assisted public health intervention for reducing engagement with alcohol consumption or smoking in a high-income country. We searched Medline (Ovid), Embase (Ovid), Web of Science (Core collection), and Scopus for relevant trials published between 2010 and 14 January 2025. We also searched for reviews of public health interventions for alcohol consumption or smoking published between 2023 and 14 January 2025, and extracted all references from relevant reviews for screening. We conducted forward and backward citation searching on all included trials (dates of searches: July to September 2025). Screening for trials was conducted independently by two reviewers. Two reviewers independently assessed risk of bias using the Cochrane Risk of Bias 2 tool. As the included trials were heterogeneous in terms of interventions, outcomes, and timepoints, we synthesised the results narratively. Results: We included 4 trials for reducing alcohol consumption (comprising 16 reports and 4,718 randomised participants): 3 trials used apps with chatbots and reported mixed evidence, and one trial, which was effective, used a rules-based AI that tailored website content. We included 12 trials for stopping smoking (comprising 31 reports and 68,659 randomised participants): 8 trials of apps or messenger chatbots and 4 trials of recommender systems, all of which had mixed evidence. For many trials, the interventions had multiple components, of which AI-assistance was only one, meaning the effectiveness of AI-assistance specifically could not be determined. Additionally, high risks of bias across almost all trials reduced confidence in the results. There were no trials using large language models (LLMs). Conclusions: There is no strong evidence of a beneficial effect of AI-assistance in public health interventions for reducing alcohol consumption or smoking in high-income countries. Future research should better describe public health interventions that use AI and embed equity considerations into their design and analysis to ensure already disadvantaged groups are not harmed further by the adoption of AI-assisted interventions in public health. Clinical Trial: PROSPERO CRD42025642317 and CRD42025642334.

  • A Web-Based Multilingual Conversational AI System for Preoperative Anaesthesia Risk Education: Randomized Equivalence Trial

    Background: Informed consent depends on patients' understanding of procedural risks, yet comprehension of anesthesia risk remains poor despite routine preoperative consultation. Current consent processes rely on time-limited clinician encounters as the primary mechanism of patient education, creating challenges for scalability, efficiency, and patient engagement. Conversational artificial intelligence (AI) offers an alternative model in which foundational education occurs before clinician contact, potentially redefining the role of the consultation itself. Objective: We evaluated whether a multilingual retrieval-augmented conversational AI system could achieve patient-reported understanding comparable to standard consultation while improving clinical workflow efficiency. Methods: We conducted a prospective randomized equivalence trial involving 130 adults undergoing elective surgery at a tertiary academic medical center. Participants were randomly assigned in a 1:1 ratio to receive either PEAR (Preoperative Education of Anaesthesia Risks), a multilingual retrieval-augmented conversational AI system grounded in institutional consent materials, followed by standard consultation, or standard consultation alone. The primary outcome was patient-reported understanding of anesthesia risk after consultation, assessed using three 5-point Likert-scale measures evaluating understanding of risks, confidence in the anesthesia plan, and ability to recall and explain key risks. Equivalence was defined a priori as a margin of ±0.5 Likert points. Secondary outcomes included pre-consultation understanding, technology acceptance, patient preference, clinical efficiency, safety, and economic impact. Results: Among 130 enrolled participants (mean age, 52.4 years; 54.6% male), post-consultation understanding in the AI group met the prespecified equivalence criterion across all primary measures. Mean between-group differences (AI minus control) were −0.03 (90% confidence interval [CI], −0.19 to 0.13), −0.05 (90% CI, −0.22 to 0.13), and 0.12 (90% CI, −0.07 to 0.32), respectively; all confidence intervals were contained within the equivalence margin. Notably, understanding scores obtained immediately after AI interaction and before clinician contact were not significantly different from post-consultation scores in the control group, suggesting that patient understanding had been established before the clinical encounter. Technology acceptance was high across all domains. Sixty-three percent of participants preferred AI-assisted education to traditional consultation. PEAR reduced combined consultation and documentation time by 19.3 minutes per patient, corresponding to an estimated annual net benefit of SGD 0.99 million (USD 0.78 million) at a single tertiary hospital. Clinicians identified documentation inaccuracies in 16.9% of encounters, all minor or moderate in severity, supporting continued clinician oversight. Conclusions: A retrieval-augmented conversational AI system achieved patient-reported understanding of anesthesia risk equivalent to standard consultation while substantially improving workflow efficiency. These findings suggest that patient education can be shifted upstream of clinician encounters, enabling consultations to focus on verification, contextualization, and shared decision-making rather than primary information delivery. Clinical Trial: The study was approved by the SingHealth Centralised Institutional Review Board (CIRB 2025/0673), registered at ClinicalTrials.gov (NCT06949462), and reported in accordance with CONSORT 2025 and CONSORT-AI Extension guidelines.

  • The SaMD Trap: A Sociotechnical Theory of Pre-deployment Failure in Regulated Digital Health

    Background: Pre-deployment failure, where digital health technologies do not progress to implementation despite technical maturity, clinical validation, and user testing, remains poorly understood and undertheorised. Objective: This study introduces the Software as a Medical Device (SaMD) Trap, a sociotechnical theory that explains how regulated digital health innovations can fail before implementation and adoption Methods: The theory was developed through a qualitative case study of a digital health tool created within a funded research programme. Data comprised 11 semi-structured interviews involving seven stakeholder groups and 17 months of programme meeting minutes. A sociotechnical analysis was conducted to identify mechanisms contributing to pre-deployment failure. Results: Three interacting mechanisms were identified. Invisible Regulatory Infrastructure reflects limited regulatory literacy and insufficient access to regulatory expertise and support. Sequential Gatekeeping describes how regulatory, governance, and compliance requirements emerge progressively and cannot be addressed independently. Temporal Misalignment occurs when regulatory and implementation timelines extend beyond the duration of research funding cycles. Together, these mechanisms explain why technically viable and clinically validated digital health systems may be unable to transition into routine practice. The theory also generated a set of testable propositions and informed the development of the SaMD Readiness Scoping Tool, designed to assess regulatory feasibility during project planning. Conclusions: The SaMD Trap provides a novel sociotechnical explanation for pre-deployment failure in regulated digital health. By identifying the mechanisms that hinder implementation before deployment begins, the theory offers a foundation for future empirical testing and highlights the importance of considering regulatory readiness alongside technical and clinical development during the design and funding of digital health innovations.

  • Multi-AI Review Enhances the Clinical Quality of AI-Generated Rehabilitation Exercise Instructional Images: A Blinded Multi-Institutional Study

    Background: AI-generated instructional images are increasingly used in rehabilitation patient education, yet single-pass generation yields anatomical inaccuracies and unsafe postures. Objective: Whether structured multi-agent AI critique pipelines measurably improve clinician-judged image quality has not been empirically evaluated. Methods: In this double-blind comparative study, 41 reviewers from three independent hospital and university teams rated 98 paired rehabilitation exercise images on a validated 14-item instrument covering four domains: clinical accuracy, instructional utility, patient safety, and clinical adoption intent. Image A was produced by single-pass AI generation; Image B underwent a four-agent critique-and-refinement pipeline. Wilcoxon signed-rank tests with Bonferroni correction, mixed-effects models, and intraclass correlation coefficients were applied. Results: Results Across 815 paired ratings, all 14 rating dimensions favored Image B after Bonferroni correction (all p < 0.001; Rank-Biserial Correlation, r_RBC 0.48–0.73), with the largest effects in visual clarity (r_RBC 0.71–0.73) and patient safety (r_RBC 0.65–0.70). At construct level, patient safety showed the greatest absolute improvement, followed by instructional utility, clinical adoption intent, and clinical accuracy. Image B was the preferred choice in 65.3% of forced-choice judgements versus 12.4% for Image A (one-sided binomial p < 0.001). Open-text defect coding showed Image B had markedly lower rates of unclear imagery (9.7% vs. 2.6%), lack of instructions (6.7% vs. 0.1%), and missing safety notes (2.5% vs. 0.4%). Quality advantages were observed across all ten anatomical regions and were independent of professional backgrounds. Conclusions: A multiple-agent critique-and-refinement pipeline substantially improves clinician-judged quality of AI-generated rehabilitation instructional images across all measured domains, with the greatest gains in patient safety and visual clarity. These findings provide the first empirical quality benchmark for pipeline-based AI image generation in rehabilitation and support multi-AI review as a scalable minimum standard before clinical deployment, pending prospective patient-outcome evaluation.

  • Artificial Intelligence Attitude Measurement Instruments in Healthcare: A Systematic Review of Measurement Properties

    Background: The rapid advancement of artificial intelligence (AI) in healthcare has led to the development of numerous instruments for assessing attitudes toward AI. Although a variety of instruments have been developed for healthcare populations, the quality of their measurement properties, the certainty of the supporting evidence, and their applicability have not been systematically evaluated. Objective: Systematically evaluate the measurement properties and methodological quality of AI attitude measurement instruments in the medical field, and to provide evidence-based recommendations for healthcare administrators in selecting appropriate measurement instruments. Methods: A systematic search was conducted in PubMed, Embase, Web of Science, and CINAHL databases to identify studies assessing attitudes toward AI among healthcare populations. The search covered all records from database inception to January 28, 2026. The methodological quality and measurement properties of included instruments were assessed following the Consensus-based Standards for the Selection of Health Measurement Instruments (COSMIN) guidelines. The quality of each instrument was rated, and overall recommendations were formulated. Results: A total of 30 studies involving 17 artificial intelligence attitude measurement instruments were included. Most instruments demonstrated satisfactory structural validity and internal consistency; however, evidence regarding content validity, cross-cultural validity, and criterion validity remained limited. Hypothesis testing for construct validity showed generally favorable results. Based on the overall assessment of measurement properties and evidence grading, nine instruments were classified as A-level recommendations, six as B-level recommendations, and two as C-level recommendations. Conclusions: AAAW demonstrated the most favorable overall measurement properties among existing artificial intelligence attitude measurement instruments in the medical field and is recommended for current use. However, the overall methodological quality of available instruments remains limited due to insufficient reporting of measurement properties and methodological procedures, as well as heterogeneity among target populations. Future studies should adhere to standardized instrument development guidelines, enhance methodological rigor and generalizability, and promote the development of reliable measurement instruments to support the evidence-based implementation of artificial intelligence in healthcare. Clinical Trial: PROSPERO CRD420261365231;https://www.crd.york.ac.uk/PROSPERO/recorddashboard

  • Background: Participants with β-thalassemia frequently experience repeated hospital visits and cannulations, leading to chronic pain and anxiety, which adversely affect their quality of life. Virtual Reality (VR) has shown efficacy in managing the pain and anxiety associated with medical procedures. The purpose of this study was to evaluate the effects of therapeutic VR on pain, anxiety, fatigue, boredom, and participant satisfaction during intravenous (IV) cannulation procedures Objective: This study evaluated the effectiveness of therapeutic VR in reducing pain and anxiety during IV cannulation among thalassemia participants, compared with SOC. Secondary objectives included assessing the impact of VR on fatigue, boredom, and participant satisfaction. Methods: A single-center, non-randomized crossover clinical trial was conducted at the Dubai Thalassemia Center, United Arab Emirates, between May and September 2024. Participants aged >7 years undergoing routine IV cannulation received SOC during their first visit, followed by VR-assisted cannulation during two subsequent visits. Outcomes were assessed using Visual Analogue Scale (VAS) scores for pain, anxiety, fatigue, and boredom. Physiological parameters, including heart rate and blood pressure, were also recorded. Comparisons between SOC and VR sessions were performed using paired statistical tests, with statistical significance set at P<.05 Results: A total of 115 participants completed at least one SOC session, and 111 participants completed at least one VR session. Overall, 82% of the participants were older than 18 years, and 51% were male. The mean anxiety score was significantly lower in the VR group (2.24 ± 2.6) than in the SOC group (2.92 ± 2.3; p=0.02). Similarly, the fatigue score was significantly lower in the VR group (1.67 ± 1.3) than in the SOC group (2.65 ± 2.2; p=0.01). Boredom scores (1.72±1.4 vs 2.67±2.2; P=.035) were also significantly reduced with VR. Pain scores were lower in the VR group but did not differ significantly from SOC (2.23±1.8 vs 2.89±2.1; P=.07). Participant satisfaction was numerically higher with VR (90% vs 85%; P=.237). No VR-related adverse events were reported, and the intervention was well tolerated. Conclusions: The usage of VR for the intervention is feasible, safe, and generally well-tolerated by thalassemia participants. VR effectively reduces anxiety and fatigue during routine intravenous cannulations. These findings are particularly relevant for participants who undergo lifelong, repeated procedures that cumulatively contribute to procedural distress. Clinical Trial: Clinical Trial Registration: https://clinicaltrials.gov/study/NCT07099196, identifier: NCT07099196, registered 2025-07-05. This trial was registered retrospectively.

  • Supporting Quality Improvement in Oral Healthcare Using Unstructured Patient Feedback: Taxonomy Development and Mixed Methods Evaluation Study

    Background: Natural language processing (NLP) is being increasing used to analyse unstructured patient feedback (UPF) for healthcare service quality improvement. Prior studies demonstrate the analytical potential of methods like topic modelling and sentiment analysis, but are largely descriptive or conceptual, with limited translation into tools for routine clinical practice. This leaves a critical knowledge gap between computational research and its realisation within (oral) healthcare service quality improvement, where large volumes of UPF are available but difficult to interpret at scale. Objective: Develop, implement, and evaluate a clinician-facing dashboard to translate unstructured patient feedback (UPF) into actionable insights for oral healthcare service quality improvement via a hybrid NLP Pipeline. Methods: We developed a four-stage hybrid NLP pipeline, combining topic modelling, sentiment analysis, and multi-label text classification. We then applied it to 57,794 reviews of NHS dental practices in England (2019–2024). Topic modelling using BERTopic identified 191 topics, with sentiment analysis via a fine-tuned DeBERTa model across four classes (positive, negative, neutral, and mixed). We iteratively consolidated topics into a ten-theme taxonomy through a hybrid approach integrating LLM-assisted classification with expert qualitative interpretation. The taxonomy informed a supervised multi-label text classifier, adapted for Google Maps Places reviews to assess portability, deploying it within a dashboard that processes real-time patient reviews, visualising thematic and sentiment insights. We evaluated the ten-theme taxonomy composition and dashboard useability through qualitative applied thematic analyses of reviews and ten semi-structured interviews with dental professionals. Results: Topic modelling generated 191 topics, consolidated into a ten-theme taxonomy. The sentiment classifier achieved F1=0.952 across four classes, while the multi-label theme classifier achieved micro-F1=0.765 and ROC-AUC of 0.935. Operationalised within a clinician-facing dashboard, these models enabled near real-time synthesis of patient feedback at practice level. Qualitative useability evaluation indicates the dashboard helps identify areas for service quality improvement that would otherwise be difficult to detect. Dual-axis representation of theme and sentiment enabled more nuanced interpretation, going beyond binary or single-label approaches. Overall, the dashboard helped summarise reviews for staff meetings and QI, but users tended to focus on negative feedback. Conclusions: Our study demonstrates a reproducible NLP pipeline to produce practical taxonomies both for oral healthcare and potentially other patient-feedback contexts. It addresses a critical knowledge gap in translating NLP research into clinician-facing tools for real-time service quality improvement. By integrating computational methods with domain-specific expert interpretation, we provide a way to bridge between data analysis and clinical application. Our findings highlight the need for hybrid approaches incorporating expert assessment to address error, bias, and useability. Meanwhile, our dashboard and mapping of its development pipeline offer a practical approach for embedding patient perspectives within routine digital healthcare to support data-driven, patient-centred service improvement in oral healthcare and beyond.

  • Background: Fatigue is a common, debilitating symptom across many chronic and post-acute conditions, yet it remains difficult to assess in routine care due to its subjective, fluctuating, and context-dependent nature. Conventional fatigue assessments rely primarily on retrospective self-report measures, which lack temporal resolution and ecological validity. Advances in digital health technologies create new opportunities to capture fatigue as it is experienced in daily life through continuous, remote, and patient-centred data collection. Objective: This study evaluates the feasibility and acceptability of a fully remote digital health platform that integrates wearable sensing, environmental monitoring, cardiorespiratory physiology, and ecological momentary assessment to capture the lived experience of fatigue in everyday life. Methods: The Understanding Patterns of Fatigue in Health and Disease study (NCT05622669) was a fully remote mixed-methods observational study. Participants with long COVID, myeloma, heart failure, and healthy controls completed either 2 or 4 weeks of monitoring. Data were collected using a wrist-worn wearable bracelet, in-home Bluetooth environmental beacons, a chest-worn ECG patch, and a smartphone application delivering ecological momentary assessments of fatigue. Data streams were integrated into unified visual representations combining activity, sleep, location, physiology, and self-reported symptoms. Feasibility was evaluated through recruitment, retention, adherence, and data completeness. Acceptability and interpretability were assessed through end-of-study interviews and optional participant feedback sessions. Results: Forty participants were enrolled and 37 completed study monitoring (retention rate 92.5%). Wearable bracelet data were available for 87% of study days, with adherence reaching 93% during periods of device operation. Ecological momentary assessments were completed on 83% of study days, whereas ECG patch data completeness averaged 72%. Twenty-two participants participated in feedback sessions. Participants reported high acceptability of the remote study procedures and considered the integrated visualisations to be plausible representations of their daily routines and symptom experiences. Contextual information derived from room-level location and environmental monitoring improved interpretation of activity and physiological data, enabling identification of behavioural patterns associated with work schedules, treatment cycles, and daily functioning that would not have been apparent from wearable-derived measures alone. Conclusions: A fully remote digital health platform integrating wearable, environmental, physiological, and self-reported data was feasible and acceptable across diverse populations experiencing fatigue. The integration of contextual information with behavioural and physiological monitoring enabled interpretable representations of daily life that participants recognised as meaningful reflections of their lived experience. These findings provide methodological guidance for future digital health studies and support the development of context-aware approaches to fatigue assessment and digital phenotyping in real-world settings. Clinical Trial: NCT05622669: Understanding Patterns of Fatigue in Health and Disease study (registered 17 November 2022)

  • Remote Digital Health Technologies for Reducing Time Toxicity in Cancer Patients: A Scoping Review

    Background: Time toxicity, defined as the cumulative burden of time spent by cancer patients during diagnosis and treatment, significantly impacts their quality of life, treatment adherence, and imposes additional stress on family caregivers. Remote digital health technologies are considered effective strategies to mitigate this burden. However, systematic evidence regarding their impact on time toxicity as a specific outcome remains fragmented, and assessment methods are not yet standardized. Objective: This scoping review aims to systematically summarize the current landscape of remote digital health technologies in reducing time toxicity among cancer patients, including their technological types, mechanisms of action, and outcome measures. The findings will provide an evidence-based foundation for clinical practice, technology development, and health policy formulation. Methods: This scoping review adhered to the Joanna Briggs Institute (JBI) scoping review framework and strictly followed the PRISMA-ScR guidelines. A comprehensive literature search was conducted across 7 English-language databases (PubMed, Web of Science, CINAHL, Embase, Scopus, Cochrane Library, IEEE Xplore) and 4 Chinese-language databases (CNKI, WanFang, VIP, SinoMed). The search spanned from database inception to April 26, 2026. Inclusion criteria comprised original research studies involving cancer patients aged ≥18 years, utilizing remote digital health technologies, and reporting time toxicity-related outcomes. Two reviewers independently screened the literature and extracted data, resolving discrepancies with a third. Quantitative data were extracted using a standardized form and descriptively synthesized. Qualitative data were processed using content analysis. Results: Fifty-nine studies were included, with most (83.1%) published since 2020. Six technology types were identified, predominantly web portals (45.8%), videoconferencing (23.7%), and WeChat (22.0%). Five mechanisms of action were found, with direct substitution (33.9%) and process compression (28.8%) being most common. Time toxicity outcomes, categorized into travel time/visits, waiting time, total outpatient/visit duration, hospital stay/postoperative recovery, and unplanned healthcare utilization, showed consistent reductions in the first two dimensions, and significant benefits in the third. However, the latter two exhibited higher heterogeneity. Analysis revealed research gaps, particularly for video conferencing in prehabilitation/remote monitoring and WeChat in remote monitoring/data-driven triage. Only 6.8% of studies applied theoretical frameworks. Qualitative findings highlighted the elimination of travel/waiting times, and downstream benefits like energy preservation and reduced caregiver burden. Conclusions: Remote digital health technologies show significant potential in reducing cancer patients' time toxicity, mainly by optimizing administrative processes and minimizing travel/waiting times through direct substitution and process compression. However, further exploration is needed for app-driven data-driven triage, wearable-assisted prehabilitation, and WeChat-based remote monitoring/triage. Future research must enhance theoretical guidance, standardize outcome measurement, and address global health equity, with nursing playing a crucial role in patient-centered digital health innovation.

  • Background: Progression is the leading cause of death among patients receiving first line treatment for diffuse large B-cell lymphoma. Some patients die from treatment-related toxicity, secondary cancers, or comorbidities, particularly cardiovascular conditions. The early detection of these events could improve quality of life and event-free survival. Objective: This study aims to describe the adverse events that have occurred and the methods used to manage them, comparing an electronic application-assisted physician approach with standard follow-up procedures. Methods: An open-label prospective, randomized, controlled phase 3 trial was conducted to compare the impact of a web-based application for monitoring events occurring during rituximab combined with cyclophosphamide, hydroxydaunorubicin (doxorubicin), oncovin (vincristine), and prednisone (R-CHOP) treatment (experimental group) with that of standard monitoring (control group). Results: A total of 62 patients were included in the study, with a median age of 62.0 years (Q1–Q3: 50.0–74.0). Thirty-one patients were randomized to each group (1:1 randomization). The study was terminated prematurely on October 17, 2024, due to bankruptcy of the unit responsible for promoting clinical trials. The median follow-up was 8.3 months (Q1–Q3: 1–24.2 months). In total, 617 events were reported (403 in the experimental group and 214 in the control group) (P<.01). About 340 events, including 238 in the experimental group and 102 in the control group, required intervention (P<.01). The median numbers of events per patient were 14 in the experimental group (Q1–Q3: 4–24), and 7 in the control group (Q1–Q3: 3–10) (P=.05). The mean times to treatment were 7.7 days in the overall population, 4.5 days in the experimental group, and 15.6 days in the control group (P=.14). Thirty-one events, including 13 in the experimental group and 18 in the control group, required hospitalization (P<.01). The event-free survival rates at the 12-months follow-up were 61.1% in the experimental group and 68.7% in the control group (P=.40). Conclusions: Event monitoring via a web application was feasible in patients with diffuse large B-cell lymphoma, with a higher number of reports in the experimental group. Compared with the control group, the experimental group had a lower proportion of events requiring hospitalization. However, considering the premature termination and the descriptive nature of the study, these results should be interpreted as exploratory Clinical Trial: ClinicalTrials.gov NCT05298293

  • Telemedicine Implementation During Prolonged Conflict: Retrospective Observational Study of Organizational Readiness

    • This manuscript needs more reviewers

    Background: Armed conflicts and prolonged security crises substantially disrupt routine ambulatory healthcare delivery. Telemedicine has increasingly been recognized as an important component of healthcare system resilience during emergencies, yet evidence regarding sustained large-scale telemedicine implementation during prolonged armed conflicts remains limited. Objective: To evaluate large-scale telemedicine implementation during a prolonged regional armed conflict and identify the organizational determinants of successful telemedicine scalability within a tertiary healthcare system. Methods: This retrospective observational study evaluated ambulatory telemedicine implementation at Sheba Medical Center, Israel's largest tertiary academic medical center, during the regional conflict beginning on February 28, 2026. Outpatient encounters performed between February and April 2026 were analyzed, including both in-person and telemedicine visits. Imaging and laboratory services were excluded because they are inherently unsuitable for telemedicine delivery. Primary analyses focused on March 2026, the only complete month during the conflict period. Telemedicine utilization, implementation patterns, and temporal trends were analyzed longitudinally. Results: Despite prolonged emergency conditions and repeated missile alerts, overall institutional ambulatory activity remained at approximately 85% of pre-conflict baseline activity during March 2026. Across the study period, 21,454 of 119,632 analyzed ambulatory encounters were conducted via telemedicine, corresponding to an overall telemedicine utilization rate of 17.9% (95% CI 17.7–18.2%), compared with 5.8% (95% CI 5.7–5.9%) before the conflict (relative risk 3.09, 95% CI 3.01–3.16; P < .001). Mean daily telemedicine activity increased from 363 to 933 encounters per day. Telemedicine expansion was heterogeneous across clinical divisions, with the greatest scalability observed in services that had integrated telemedicine into routine practice before the conflict. Large-scale implementation was supported by same-day conversion of scheduled visits, clinician enablement, centralized operational oversight, and continuous monitoring of ambulatory activity. Conclusions: Large-scale telemedicine implementation was associated with sustained ambulatory care delivery during a prolonged armed conflict and functioned as a key component of a hybrid continuity-preservation strategy. Successful telemedicine scalability depended less on technology itself than on organizational readiness established before the crisis, including pre-existing clinical integration, operational governance, institutional adaptability, and clinician familiarity with telemedicine delivery. These findings suggest that healthcare systems seeking to strengthen resilience should integrate telemedicine into routine clinical practice before emergencies occur rather than relying on crisis-driven implementation.

  • Background: Large language model-based intelligent standardized patient systems offer a scalable solution for clinical skills training, yet rigorous comparisons against active pedagogical controls remain scarce, and the mechanisms underlying their effectiveness are poorly specified. This study evaluated an LLM-based intelligent SP system compared with small-group, tutor-facilitated case discussion in dental education, with the explicit objective of disentangling the role of differential individual active learning time as a potential mediator of observed effects. Objective: To evaluate an LLM-based intelligent standardized patient system compared with small-group, tutor-facilitated case discussion in dental education, and to disentangle the role of differential individual active learning time as a potential mediator of observed effects. Methods: In this single-center, parallel, open-label randomized trial conducted from May to June 2025, 60 third- and fourth-year dental students were randomized 1:1 to the AI-SP group (n=30) or the TC-SG group (n=30). The AI-SP group completed six interactive virtual patient modules powered by DeepSeek-V3 with automated feedback, while the TC-SG group discussed identical cases in groups of four to five students with a tutor. The primary outcome was the adjusted post-intervention mini-CEX Overall Score assessed by a blinded expert panel using ANCOVA with baseline adjustment. Secondary outcomes included AI-generated communication metrics and clinical self-efficacy. The study design inherently produced substantially different individual active learning time between arms—approximately 110–120 minutes per session for AI-SP versus 24–30 minutes for TC-SG—which we explicitly treated as a design-defined dose parameter. Due to significant baseline imbalances in secondary AI-generated metrics favouring the TC-SG group, causal inferences were restricted to the primary outcome, where baseline balance was confirmed. Results: Fifty-nine participants completed the trial (AI-SP=30, TC-SG=29). For the primary outcome, the AI-SP group showed significantly higher adjusted post-intervention mini-CEX Overall Score compared with TC-SG (adjusted mean difference [aMD]=0.68; 95% CI, 0.29–1.07; P=0.001; η²p=0.22). This effect size corresponded to the substantial difference in individual practice density between the two conditions. For AI-generated secondary metrics, despite significant within-group improvements in the AI-SP group (all P<0.001), no significant between-group differences were observed at post-intervention for Accuracy (P=0.059) or Interactivity (P=0.161), indicating comparable endpoint performance. Given the baseline imbalance in these metrics, we interpret these null between-group differences conservatively, without claims of catch-up or superiority. An exploratory regression analysis treating estimated individual active learning time as a continuous predictor revealed a significant association with mini-CEX improvement (β=0.34, P=0.002), supporting the dose-response interpretation. Conclusions: The AI-SP system, by affording substantially higher individual active learning density, produced superior expert-assessed clinical outcomes compared with small-group discussion. However, this advantage is confounded by unequal practice time and cannot be attributed solely to AI intelligence. Our findings suggest that AI-SP functions primarily as an effective pedagogical platform for scaling individual deliberate practice opportunities. Future research must employ dose-equated designs to isolate the unique contribution of AI-generated feedback from the general benefits of increased active learning time.

  • Remote Patient Monitoring in Orthopaedic Surgery: Leveraging Digital Health for Enhanced Patient Care

    Remote patient monitoring (RPM) includes a variety of technologies for evaluation of patient symptoms and joint motion beyond conventional clinical settings, with the goals of increasing access to care and possibly decreasing healthcare costs. RPM utilizes modern app-based and sensor-based technology to transmit qualitative and quantitative data to orthopaedic surgeons and their assistants for analysis and possible intervention. RPM technology includes smart phone mobile applications, telemedicine, portable wearable motion sensor devices, and a knee arthroplasty component. RPM in orthopedic surgery may improve collection of patient-reported outcomes (PROMS) and may increase physician reimbursement with new billing codes. The data collected could help individualize perioperative care, increase access to care in rural areas, and possibly empower patients to take a more active role in their care. RPM could possibly decrease healthcare costs with a reduction of unnecessary hospital readmissions or emergency room visits. It may also be used for preventive and post-procedural services. As RPM for orthopedic surgery patients may have more widespread use in the future, orthopedic surgeons should understand its use for providing musculoskeletal care and the appropriate billing codes for reimbursement.

  • Test-Retest Reliability of Smartwatch-Derived Features for Longitudinal Monitoring: An Observational Cohort Study

    Background: Smartwatches are increasingly used for decentralized data collection in clinical research, but the everyday-life settings that make these data attractive also introduce variability. Before a smartwatch-derived feature can support clinical monitoring, its reproducibility must be established: features with low test-retest reliability weaken associations with clinical outcomes and potentially generate non-actionable signals. Reliability is expected to vary by feature type, aggregation window, and data availability, but has not been systematically screened in a clinical cohort. Objective: This study aimed to evaluate the test-retest reliability of candidate smartwatch-derived features for longitudinal monitoring in adults with advanced cancer and in healthy controls, and to determine how reliability depends on the temporal aggregation window. Methods: In a prospective single-centre observational cohort study, we analysed 8 weeks of Garmin Vivosmart 5 sensor data from 60 adults with advanced cancer, and 20 healthy controls. Out of 80 participants, 77 contributed usable smartwatch data. We examined 35 daily features across seven domains: heart rate variability, heart rate, respiration, oxygen saturation, sleep, activity, and smartwatch-derived stress. Test-retest reliability was quantified as the intraclass correlation coefficient (ICC(2,1)) across adjacent non-overlapping 1-, 3-, and 7-day windows, with 95% CIs from subject-level bootstrap resampling (10,000 resamples). Between-group and therapy-centred contrasts used permutation testing with Benjamini-Hochberg false discovery rate correction. Results: Reliability improved with longer aggregation windows in both cohorts. Between 1-day and 7-day windows, median ICC(2,1) increased from 0.53 to 0.73 in controls and from 0.66 to 0.79 in patients. The number of features reaching good-to-excellent reliability, defined as ICC(2,1)≥0.75, increased from 5 of 35 to 15 of 34 in controls and from 12 of 35 to 26 of 35 in patients. Heart rate and heart rate variability features were the most reliable, with 4 of 11 reaching weekly ICC(2,1)≥0.90 in both cohorts. Activity, respiration, sleep and oxygen-saturation features were more sensitive to aggregation window and data availability, showing larger gains from daily to weekly aggregation (e.g., step count ICC increased from 0.30 to 0.73 in controls). Weekly reliability did not differ significantly between patients and controls (median ΔICC=0.036, no feature survived FDR correction, all q>0.05). No feature showed a significant change in reliability around therapy (median ΔICC=0.016, all q>0.05). Conclusions: Weekly aggregation improved the reliability of many smartwatch-derived features, but reliability remained feature specific. Heart rate and heart rate-variability features were consistently reliable, whereas selected sleep and oxygen-saturation features displayed only moderate reliability across all aggregation windows. Reliability was comparable across cohorts and stable around therapy, indicating that feature-wise estimates are transferable across these clinical contexts. Feature-level reliability screening is a prerequisite before smartwatch-derived measures are used in clinical monitoring.

  • Digital Technology–Supported Exercise Rehabilitation for Knee Osteoarthritis: A Systematic Review and Three-Level Meta-Analysis of Randomized Controlled Trials

    • This manuscript needs more reviewers

    Background: Background: Exercise rehabilitation is a core component of knee osteoarthritis management, but sustained delivery may be constrained by time, geography, and limited health care resources. Digital technologies may improve access to prescribed exercise and support its delivery and monitoring; however, their overall effectiveness and the factors influencing treatment response remain uncertain. Objective: Objective:To systematically evaluate the effects of digital technology–supported exercise rehabilitation on pain, patient-reported function, performance-based physical function, and quality of life in people with knee osteoarthritis, and to investigate potential moderators of treatment effects. Methods: Methods: PubMed, Web of Science, Embase, EBSCO, the Cochrane Library, China National Knowledge Infrastructure, and Wanfang Data were systematically searched from inception to May 20, 2026. Randomized controlled trials were eligible if they evaluated exercise-based rehabilitation supported by digital technologies, including telerehabilitation, mobile applications, or virtual reality. Three-level random-effects models were used to account for dependencies among multiple effect sizes arising from different outcomes, measurement instruments, or assessment time points within the same study. Statistical inference was further supported by cluster-robust variance estimation with CR2 small-sample correction. Subgroup analyses and meta-regression were conducted to examine the potential moderating effects of knee osteoarthritis severity, type of digital technology, comparator condition, sensor use, intervention duration, and weekly intervention frequency. The certainty of evidence was assessed using the GRADE approach. Results: Results: Twenty randomized controlled trials involving 1,222 participants with knee osteoarthritis were included. Three-level meta-analyses showed that digital technology–supported exercise rehabilitation significantly reduced pain (Hedges’ g = 0.60, 95% CI [0.15, 1.06], 95% PI [−1.29, 2.50]) and improved patient-reported function (Hedges’ g = 0.72, 95% CI [0.24, 1.20], 95% PI [−1.42, 2.87]), performance-based physical function (Hedges’ g = 0.54, 95% CI [0.11, 0.97], 95% PI [−0.99, 2.07]), and quality of life (Hedges’ g = 0.51, 95% CI [0.16, 0.85], 95% PI [−0.46, 1.48]). These effects remained statistically significant after CR2 small-sample correction. Subgroup analyses indicated greater improvements in quality of life among participants with mild-to-moderate knee osteoarthritis. Meta-regression further showed that longer intervention duration was associated with greater pain reduction and improvements in quality of life, whereas a higher weekly intervention frequency was associated with greater improvements in patient-reported function. Sensitivity analyses supported the robustness of the findings for patient-reported function and quality of life. Although the pooled estimates consistently favored digital technology–supported exercise rehabilitation, all 95% PI crossed the line of no effect, indicating substantial variability in the treatment effects that may be observed across future clinical settings. Conclusions: Conclusions: Digital technology–supported exercise rehabilitation may improve pain, function, and quality of life in people with knee osteoarthritis. Its principal clinical value may lie in expanding access to exercise rehabilitation and supporting the implementation of prescribed exercise. however, the optimal mode of delivery remains to be established through further high-quality research. Clinical Trial: Trial Registration: PROSPERO CRD420261404719; https://www.crd.york.ac.uk/PROSPERO/view/CRD420261404719

  • Background: Digital health has the potential to improve maternal and child health (MCH), particularly during the first 1,000 days of life. However, its successful implementation and adoption depend not only on the availability of technology but also on organizational readiness, governance, and the broader sociotechnical context. Evidence on how these factors interact in low- and low-middle income settings remains limited. Objective: This study examined the interplay of sociodemographic context, facility governance, and ICT readiness in shaping the adoption of a digital health tool for MCH in Manila, Philippines. Methods: A qualitative descriptive study underpinned by a constructivist paradigm was conducted in Manila, Philippines. The study sites included two outpatient departments in the Philippine General Hospital (PGH) and government primary healthcare facilities. Participants were selected through purposive sampling. Data were collected through a desk review, facility observations, and semi-structured in-depth interviews and focus group discussions. Qualitative data were analyzed using a two-phase approach combining rapid qualitative analysis and reflexive thematic analysis, guided by the Network of Influence Framework. Results: Readiness for digital health adoption was shaped by the interaction of favorable user-level factors and persistent organizational and governance constraints. Manila demonstrated high levels of digital engagement, with widespread smartphone ownership and internet access among potential users. However, fragmented health information systems and limited interoperability continue to constrain digital transformation across healthcare facilities. Six themes emerged from the qualitative analysis. Barriers included limited internet connectivity, inadequate availability of digital devices, usability challenges, and the lack of localized and integrated digital health tools. Enablers included interoperability across health information systems and positive attitudes toward technology adoption among mothers and healthcare providers. Although participants expressed strong willingness to adopt digital health innovations, structural and governance constraints limited their routine implementation. Overall, the findings indicate that the study sites currently exhibit more barriers than enablers, reflecting an imbalance in readiness to support the implementation of the Kalinga application. These findings provide important considerations for prioritizing context-responsive features and developing implementation strategies to address identified readiness gaps. Conclusions: Despite high levels of digital engagement among potential end users, structural and governance factors continue to impede digital health adoption for MCH in Manila. Bridging the gap between user readiness and system capacity requires investments in digital infrastructure, interoperability, localized digital solutions, and integrated health information systems.

  • Background: Childhood cancer remains the leading cause of disease-related death among children and adolescents worldwide, despite increasing survival rates. Survivors often face long-term physical and psychological side effects. In response, mobile health (mHealth) has emerged as a promising tool to support pediatric oncology patients. However, adoption and sustained use of mHealth interventions vary, often due to usability, accessibility, acceptability, feasibility, and user satisfaction challenges. Objective: This systematic review aims to synthesize existing literature on the usability, accessibility, acceptability, feasibility, user satisfaction, and overall user experience of mHealth interventions in pediatric oncology. Methods: This systematic review was conducted according to ENTREQ and PRISMA guidelines. Studies were identified through a comprehensive database search (Medline, Embase, Web of Science, Cochrane CENTRAL, and Google Scholar) performed by a medical information specialist. Screening and selection were independently performed by three reviewers using Rayyan. Data extraction included intervention characteristics, participant demographics, and reported outcomes. Thematic analysis was used to synthesize the reported outcomes across the included studies. The methodological quality of the included studies was assessed using the Critical Appraisal Skills Programme (CASP) checklist. Results: Of 12,620 studies identified, 13 were included in this systematic review. Thematic analysis of the included studies revealed four themes, each encompassing multiple subthemes. These were: Empowerment and Participation in Care, Engagement Through Design and Motivation, Usability and User Experience, Informational Support and Peer Connection, and System-Level Limitations and Disconnects. Children and adolescents were found to play an active role in their care, using mHealth applications to track symptoms, support self-management, and enhance adherence. Design and personalization features—such as narrative elements, gamification, and visual appeal—played a key role in sustaining engagement. Usability was influenced by factors including digital literacy, clarity of instructions, and intuitive navigation, with notable differences across age groups. mHealth tools were also valued for their capacity to deliver trusted, peer-mediated information and foster meaningful connections with others facing similar health journeys. Conclusions: mHealth interventions show a promising role in supporting pediatric cancer care, particularly by enhancing engagement, empowerment, and overall user experience. However, their effectiveness depends on user-centered, developmentally appropriate design and more diverse, long-term research to ensure they align with the real-world needs of pediatric oncology patients.

  • A scoping review of dashboards to measure and improve quality of care in oncology

    Background: Quality dashboards are increasingly being used in oncology to support monitoring and standardisation of quality of care against defined benchmarks, yet there is limited consensus or guidelines on effective dashboard design. Objective: This scoping review aims to summarise the evidence supporting the use of quality dashboards in improving cancer care and to identify key design features to inform future development. Methods: A comprehensive literature search of MEDLINE (PubMed) and EMBASE (Ovid) was conducted on June 12, 2026, using keywords including “performance,” “cancer,” “quality,” and “dashboard.” Reference lists of included studies were screened, and updates to previously published studies were searched to identify additional relevant articles. Eligible studies were full-text English-language publications involving adult patients undergoing work-up, treatment, or follow-up for solid organ malignancies in the outpatient setting that described the design and utilisation of digital dashboards capturing data on quality indicators. Results: Of 181 abstracts retrieved, 16 met inclusion criteria. An additional 18 papers were identified through reference list screening and searches for updated publications, resulting in 34 included studies. Across a range of cancer types, there was marked heterogeneity in the quality indicators, outcome measures, and reporting frequency. Despite this, there is evidence that implementing quality dashboards can improve cancer care processes, outcomes, and adherence to best practice guidelines. However, notable underexplored areas include the impact on recurrent disease management, end-of-life care, and survival outcomes. Effective design features include a single overview page with more information provided on click-through, colour coding to indicate performance levels, and funnel plots to identify outliers. Conclusions: Quality dashboards can improve guideline-recommended care processes across various stages of cancer care, with considerable scope for improvement, including through standardisation. Further research is needed to assess their impact in underexplored but clinically important areas, such as recurrent disease or end-of-life care settings. Clinical Trial: N/A

  • Using Design Thinking and Co-Design Methods to Align Diverse User Needs in Care Transitions: Designing the Digital Bridge.

    • This manuscript needs more reviewers

    Background: Developing digital health solutions for complex health service environments poses a unique design challenge. Service models, like transitions from hospital-to-home for older adults with complex care needs, represent dynamic contexts in which multiple user types are working across diverse settings and workflows. Designing digital health solutions for these types of environments requires advancing traditional methods intended to design for the needs of single user groups working in more bounded settings. Objective: This study addresses this design challenge by combing Design Thinking and co-design methods to develop a digitally enabled hospital-to-home communication platform to meet the needs of older adult patients with complex care needs, their caregivers and hospital and primary care clinicians who are involved in the transition process. To co-design the Digital Bridge tool, the study was guided by two questions addressed in this paper: 1) How can we translate diverse user groups’ needs into technology features? And 2) Can incorporating a Design Thinking-driven process as part of user-centred co-design help to manage tensions with diverse user groups perspectives? Methods: The Institute of Design at Stanford’s Design Thinking approach (empathize, define, ideate, prototype and test) was applied to guide co-design of the tool; leveraging multiple virtual platforms (e.g. Zoom, Jamboards, journey maps), research methods (e.g. interviews, working groups, asynchronous feedback, surveys), and informatics tools (e.g. information flow and business process maps). Working groups consisting of patients and caregivers, hospital and primary care clinicians, and the project’s Citizen Advisory Committee. Working groups engaged in multiple-iterative design phases with the research and design team to adapt two existing technologies into the new Digital Bridge platform. Results: Pre-design work involving interviews with patients and caregivers identified nine challenges in the hospital-to-home transition process, including communication barriers, feeling rushed and invisible, and not knowing where to go when help was needed. Challenges acted as design anchors, guiding a series of five working group sessions with patients and caregivers (n=9), five working group sessions with hospital (acute and rehab sites) and primary care clinicians (n=28) and a round of surveys. Needs were mapped onto six tangible design functions and integrated in two separate technology platforms to align to the local digital ecosystems and workflows at two hospitals networks in Canada. Working across user groups, tensions around language and workflow surfaced and were addressed by prioritizing the patient- and caregiver-identified challenges, while maintaining a person-centred lens. Conclusions: The Design Thinking approach was useful in guiding the co-design process, however, effective collaboration across diverse user groups required iterative rather than linear movement between stages of the Design Thinking approach. Future co-design projects working across diverse groups should consider embedding shared-empathy activities to manage tensions and develop true person-centred solutions. Clinical Trial: ClinicalTrials.gov NCT04287192; https://clinicaltrials.gov/ct2/show/NCT04287192

  • Background: Excessive gestational weight gain (GWG) is associated with a range of adverse maternal and neonatal outcomes. Digital health interventions (DHIs) are increasingly used in prenatal care, but their effectiveness for GWG management remains uncertain, and the factors that may influence intervention effects have not been fully clarified. Objective: This systematic review and meta-analysis aimed to examine the effectiveness of DHIs on GWG outcomes during pregnancy. Methods: We conducted a systematic review and meta-analysis of randomized controlled trials evaluating DHIs for GWG management during pregnancy. PubMed, Embase, Medline, Cochrane CENTRAL, and ClinicalTrials.gov were searched from inception to January 16, 2026. Primary outcomes included total GWG, weekly GWG, and excessive GWG according to Institute of Medicine recommendations. Random-effects models were used to pool mean differences (MDs) and risk ratios (RRs) with 95% confidence intervals (CIs). Prespecified subgroup analyses explored potential effect modifiers, including pre-pregnancy body mass index (BMI), gestational age at intervention initiation, geographic region, and gestational diabetes mellitus (GDM) status. Results: Forty randomized controlled trials involving 8,178 pregnant women were included. Compared with usual care, DHIs significantly reduced total GWG (MD = −0.68 kg, 95%CI = −1.11 to −0.25) and weekly GWG (MD = −0.05 kg/week, 95%CI = −0.09 to −0.02). DHIs also reduced the risk of excessive GWG (RR = 0.85, 95%CI = 0.77 to 0.94). Subgroup analyses showed that DHIs reduced total GWG by −1.34 kg (95%CI = −1.69 to −1.02) among women with pre-pregnancy overweight or obesity, whereas no significant reduction was observed among mixed BMI women. Regarding intervention timing, HDIs initiated at or before 20 weeks’ gestation yielded significant reductions in weekly GWG and risk of excessive GWG. Conversely, initiating DHIs after 20 weeks showed greater reductions in total GWG. Geographically, DHIs significantly lowered total GWG in Europe and Asia, but no such effect was observed in North America. Meta-regression analyses did not identify significant linear associations between pre-pregnancy BMI and GWG outcomes. Sensitivity analyses supported the robustness of the findings, although publication bias was detected. Conclusions: DHIs can achieve modest but significant improvements in GWG management during pregnancy and may help reduce the likelihood of excessive GWG. The benefits appear more evident among women with pre-pregnancy overweight or obesity and in interventions initiated earlier in pregnancy. Given their accessibility and scalability, DHIs may represent a useful complement to routine prenatal care, although further high-quality large-scale trials are still needed to refine intervention strategies and determine the most effective implementation approaches.

  • Background: Multicomponent digital health interventions are increasingly used to support type 1 diabetes (T1D) self-management and are generally acceptable to patients. However, existing evaluations primarily report average effects at the group level, with limited understanding of how outcomes arise across individuals, intervention components, and engagement patterns. Objective: To identify the contextual profiles under which favourable self-efficacy was observed within a multicomponent digital health intervention for emerging adults living with T1D. Methods: We conducted an embedded evaluation using Coincidence Analysis, a configurational method that examines how combinations of factors are associated with outcomes in complex interventions. The analysis was embedded within a randomised controlled trial of Keeping in Touch (KiT), a 12-month personalised text message-based intervention for emerging adults living with T1D in Canada. Self-efficacy was dichotomised as favourable or unfavourable based on baseline level and change over 12 months. We first examined contextual profiles associated with outcomes in the full sample, including the role of overall intervention exposure within these profiles. Among intervention-group participants, we further examined how patterns of exposure to intervention components and participant engagement were associated with outcomes. Results: Among 168 participants in the full-sample analysis and 89 intervention-group participants in the engagement analysis, multiple distinct profiles were associated with favourable and unfavourable outcomes. Exposure to the intervention was associated with favourable outcomes only within a specific profile characterized by shorter diabetes duration (<10 years) and lower baseline HbA1c (<9.0%). Higher engagement, characterised by a higher prompt response rate (≥Q2) and use of optional features, featured in favourable profiles, whereas lower engagement featured in unfavourable profiles. However, engagement alone was neither necessary nor sufficient for benefit. Instead, its association with outcomes depended on personal context, including diabetes duration, sex, and exposure to specific educational topics. Conclusions: For emerging adults with T1D, a one-size-fits-all approach may be insufficient, and evaluations focus solely on average intervention effect may obscure important differences in effectiveness across subgroups. By identifying the multiple context-dependent pathways associated with both outcomes, the findings underscore the need for context‑responsive, person‑centred approaches that tailor intervention content and engagement strategies to diverse population.

  • Patients’ Expectations of Physician Use of Artificial Intelligence: Systematic Review

    Background: As artificial intelligence (AI) is increasingly integrated into clinical workflows, traditional models of the patient-physician relationship are being redefined. Understanding how AI shifts patient expectations of physicians is critical for maintaining trust, ensuring accountability, and guiding medical education. Objective: To systematically review and synthesize empirical evidence on patients’ expectations of physicians who use AI in clinical decision-making. Methods: A systematic search was conducted in PubMed and Web of Science for empirical studies published between January 1, 2000, and January 31, 2026. Eligible studies examined triadic patient-physician-AI contexts and reported on patient perspectives regarding physician roles, competence, communication, or responsibility. Data were synthesized using thematic analysis and mapped onto established theoretical perspectives, including role, trust, attribution, and agency based-perspectives. Results: A total of 16 studies met the inclusion criteria, spanning diverse clinical contexts such as oncology, radiology, and primary care. Mapping findings to role-based perspectives, patients viewed clinical judgment as non-delegable (Theme 1), expecting physicians to maintain diagnostic oversight and take dual responsibility for algorithmic outputs. Final accountability remained anchored in the clinician (Theme 3), as patients continue to place the moral and legal responsibility for medical outcomes within human agency. AI literacy was also recognized as an emerging component of professionalism (Theme 7). From trust-based perspectives, confidence in AI was mediated and context-dependent (Theme 4), often grounded in the existing patient-physician relationship, provided that humanistic care was preserved (Theme 5). Patients also identified ethical responsibilities, including regulatory approval of AI use, data privacy, informed consent, patient safety, and fairness (Theme 6), as non-transferable duties of physicians. From agency-based perspectives, explaining AI was viewed as a fundamental clinical responsibility (Theme 2), requiring physicians to interpret complex AI-generated outputs for patients. Despite its perceived benefits, AI was consistently positioned as a supportive third-party in decision-making (Theme 8), with physicians expected to act as the primary mediator who contextualizes technical evidence within individual value systems. Furthermore, physicians were expected to preserve patient autonomy and facilitate shared decision-making, advancing the core principles of patient-centered care (Theme 9). Although not a primary perspective, attribution-based considerations also shaped expectations regarding physicians’ professional competence, informed consent, and patient autonomy (Theme 1, 6, and 9). Conclusions: In the context of AI-assisted care, patients articulate expanded expectations of physician responsibility, encompassing non-delegable clinical judgment, ultimate accountability, and the preservation of humanistic care. Medical education and health system governance must therefore prioritize the cultivation of augmented professionalism, ensuring that AI integration enhances rather than undermines the relational core of the physician-patient relationship. Clinical Trial: PROSPERO CRD420261295341

  • Background: Digital technology can improve diabetes treatment and management. However, the effectiveness of using digital education interventions in improving the glycaemic control of children and young people is unknown. Objective: To explore the evidence-based literature on the effectiveness of digital educational interventions in children and young people living with diabetes. The review aimed to identify online resources and technology and synthesise the effect size of interventions on glycated haemoglobin (HbA1c), in addition to other outcome measures used to assess the efficacy of the intervention. Methods: A systematic review and meta-analysis were conducted using the Joanna Briggs Institute (JBI) Methodology. A database search was completed using MEDLINE, CINAHL, Cochrane Library, Embase, ClinicalTrials.gov website, the International Clinical Trials Registry Platform, and ProQuest Dissertations. Only studies published in English and published during the last 20 years were included. An a priori protocol was developed and made available on the Open Science Framework and was registered in PROSPERO (CRD42024599125). Results: A total of 14 studies, comprising 1330 participants from 9 countries, were included. A statistically significant reduction in HbA1c levels in children and young people diagnosed with type 1 diabetes was found (MD= ꟷ0.17, 95% CIꟷ 0.29,ꟷ0.05, P = 0.006, I2 = 38%). The use of telemedicine platforms, including the transmission of blood glucose data with feedback or wearable devices, was the most common platform and form of data collection. In the subgroup analysis, the fixed effects model showed positive outcomes for diabetes related worry (MD = 2.59, 95% CI 0.77, 4.42, P = .005, I2 = 87%) and treatment satisfaction (MD= 1.92, 95% CI 0.78, 3.05, P = .001, I2 = 0%), favouring the use of 'digital educational interventions. Subgroup analyses included the duration of interventions as well as the types and content of digital educational interventions. However, the effect of these interventions according to age group and digital platform remains uncertain. Conclusions: Engagement with interventions using 'digital educational interventions' demonstrated improved HbA1c levels in children and young people with diabetes. Since we were unable to locate studies among children and young people with type 2 diabetes or prediabetes, our findings are limited to type 1 diabetes. Establishing guidelines for the design of digitally interactive interventions informed by motivational theory, the inclusion of longer follow-up times, the inclusion of low- and middle-income countries and the development of interventions for culturally and linguistically diverse populations would improve study quality, consistency of reporting and development in this emerging field.

  • Theory-Guided Development and Usability Evaluation of a Web-Based Self-Management Module for Older Adults with COPD: A Mixed-Methods Study

    Background: Chronic obstructive pulmonary disease (COPD) imposes a substantial burden on older adults, yet existing digital self-management interventions often fail to address age-specific usability barriers and lack integration of established behavioral and technology acceptance theories. While web-based platforms hold promise for supporting COPD management, few have been systematically developed with direct input from older patients and validated through rigorous mixed-methods usability evaluation in the Chinese healthcare context. Objective: To describe the theory-guided development process and evaluate the usability of a web-based self-management module embedded within the SLH-COPD platform, specifically designed for older adults with COPD in China. Methods: This study employed a sequential exploratory mixed-methods design comprising three phases. In Phase 1 (Module Development), the COPD digital health intervention module was developed and finalized based on findings from prior research, a systematic literature review, and two rounds of Delphi expert consultation. Phase 2 involved the technical configuration and integration of the module into the SLH-COPD platform. In Phase 3 (Validation), alpha and beta testing were conducted with older adults with COPD; usability was assessed using the UMUX alongside objective behavioral data to inform iterative refinement of the module. Results: Using a mixed-methods design, this study successfully constructed and optimized a web-based self-management module for older adults with COPD, embedded within the SLH-COPD platform. A cross-sectional survey (n=199) identified five core domains of user needs, including symptom monitoring, medication management, and rehabilitation exercise. Following two rounds of Delphi Method (n=17, authority coefficient Cr=0.88), expert consensus was satisfactory, with Kendall’ s W values of 0.230, 0.321, and 0.285 for first-, second-, and third-level indicators, respectively (P<0.05). The final intervention framework comprised 5 first-level, 14 second-level, and 34 third-level indicators, demonstrating excellent content validity (S-CVI/Ave=0.988) and item-level CVIs (I-CVI) ranging from 0.80-1.00. Consistency testing using the AHP yielded a random CR of 0.009 (<0.1), indicating a scientifically sound weight allocation. Alpha testing resolved technical issues such as medication reminder delays and insufficient Shaanxi-localized content. Subsequent Beta testing revealed a mean UMUX score of 77.25 (SD=2.06), significantly exceeding the accepted usability threshold, with a task completion rate > 85%. Qualitative feedback confirmed that senior-friendly designs and plain-language, localized content effectively improved the user experience among low-literacy older adults. Conclusions: This study confirms that the theory-guided web-based self-management module for older COPD patients demonstrates adequate content validity and usability, with senior-friendly design effectively meeting user needs. It provides a replicable development and validation paradigm for digital chronic disease interventions in older adults. Future randomized controlled trials are needed to evaluate its clinical effectiveness and long-term adherence. Clinical Trial: Chinese Clinical Trial Registry (ChiCTR number):PID331832

  • Trauma-Informed Language as a Safety Standard for AI and Digital Health: Lessons From Intimate Partner Violence Survivor Support

    Technology-mediated services, including chat platforms, social media, mobile applications, and emerging artificial intelligence (AI) tools, are increasingly used to support survivors of intimate partner violence (IPV). These tools can expand access to information and support, particularly for survivors who face barriers to in-person services, such as a partner’s controlling behaviors, geographic distance, transportation, childcare, or concerns about privacy and safety. However, safety in technology-mediated services is not limited to protecting survivors’ privacy and collected data. It also depends on how technologies communicate with survivors. Although risks related to privacy and confidentiality are widely recognized, this Viewpoint draws attention to an underrecognized safety concern: the potential for language used or generated by technology to cause distress, reinforce bias, stereotype, and stigma, or re-traumatize survivors. Language is not neutral. It reflects dominant social norms, power structures, and the perspectives of those with greater access to power, privilege, and resources. As a result, even language that appears respectful or objective may carry bias, stereotypes, victim-blaming narratives, or assumptions that marginalize IPV survivors. Explicitly discriminatory, bigoted, or hateful language may be more readily recognized. More difficult to identify is language that appears neutral but minimizes survivors’ concerns, misinterprets their experiences/thoughts/feelings, implies responsibility for the violence they experienced, or excludes the experiences of male, nonbinary, disabled, racialized, immigrant, or otherwise marginalized survivors. These risks are heightened in digital interactions that rely primarily on written communication because they lack tone, facial expression, body language, and other contextual cues. This concern applies across technology-mediated services but becomes especially urgent with generative AI. Because AI systems are trained on large bodies of historical language data, they may reproduce and amplify entrenched social inequities. Without intentional trauma-informed design and evaluation, AI-generated responses may threaten survivors’ perceived safety and trust in technology, disempower them, and potentially discourage future help-seeking. Drawing on the six principles of a trauma-informed approach, this Viewpoint introduces trauma-informed language as communication that recognizes the widespread impact of trauma, acknowledges that language itself can cause or exacerbate harm, and actively resists re-traumatization through language that promotes safety, trustworthiness and transparency, support, collaboration, empowerment, and attention to cultural, historical, and gendered contexts. This perspective shifts the field from reactive approaches that detect and mitigate harmful outputs after they occur toward proactive prevention. It also reframes language not as a stylistic concern, but as a core design, safety, and equity standard for digital technologies. Future work should develop and test trauma-informed language frameworks, dictionaries, and evaluation criteria across diverse survivor populations and technology contexts to ensure that digital innovation advances not only access, but also safety, dignity, equity, and healing.

  • Should Clinical Foundation Models Reason Through Disease Labels? A Falsifiable Case for Diagnosis as an Interface

    Clinical foundation models are increasingly trained on longitudinal electronic health records, learning continuous, high-dimensional patient representations that are not organized around the diagnostic vocabulary. Yet these systems are still built, evaluated, and governed as if the disease label were the natural unit of machine reasoning: the diagnosis is the privileged prediction target and the unit in which the model is expected to reason. In this Viewpoint we argue that this inherited assumption should be reversed. A disease label is a compressed, human-compatible abstraction whose usable resolution was bounded not by biology alone but by what clinicians and institutions could reliably name, teach, remember, and share. Foundation models relax that constraint, because the representation used to reason need no longer be human-readable: a machine can reason over a higher-dimensional latent patient state and render a named diagnosis only when a clinician, payer, regulator, or registry requires one. We therefore separate three things the label conflates—the internal representation a model reasons over, the clinical decision it is optimized against, and the human- and institution-facing code it renders—and reframe diagnosis as an external interface, a projection from that internal representation into a human-compatible code, rather than the substrate of machine reasoning. The claim is empirical, not rhetorical, and we hold it to a falsifiable test: comparing label-based against foundation-model latent representations, under matched data and compute, on outcomes defined outside the diagnostic coding system—treatment response, trajectory, dose, timing, and toxicity. We specify one such test in BCR–ABL-positive chronic myeloid leukaemia. The boundary condition is explicit: where a label is already a sufficient statistic for the decision, a richer representation buys nothing.

  • Background: Estimating the magnitude of effect of immunosuppressive therapy in autoimmune disease and transplantation medicine as a single numeric score in tabular data is highly valuable, especially when developing statistical or machine learning models for prognosis, risk stratification, and other tasks. Objective: We aimed to derive a single continuous score that represents a patient’s overall immune status at a given time after exposure to immunosuppressive therapy. Methods: We developed an immunosuppressive intensity (ISI) score model to estimate point-in-time immunosuppressive state as a continuous cumulative score ranging from 0 to 1. Model structure and parameterization were informed by a structured expert elicitation process using a modified Delphi approach across commonly used immunosuppressive therapies. A base ISI model was implemented as a scaled and shifted sigmoidal ISI score function incorporating three parameters: A (starting intensity), n (decay rate), and d (half-life, 50% pharmacodynamic effect). Age and lymphocyte/CD19 B-cell counts were then incorporated to generate a biomarker-informed adjusted ISI score. Finally, we developed a cumulative ISI score to model the immunosuppression state when multiple medications are active contemporaneously. We evaluated the model in three international ANCA-associated vasculitis cohorts (RITA Ireland vasculitis (RIV) registry, IDIBELL registry, and Chapel Hill registry) for biological alignment, clinical plausibility across disease phases, and utility in relapse-risk modeling compared with conventional categorical treatment encoding, using a generalized estimating equation (GEE) model. The biological alignment was further evaluated by correlating the ISI score with Torque Teno virus (TTV) count, a marker of immunosuppression. Results: Following a Delphi process, we defined parameters for the base and adjusted the ISI score across a range of intravenous and continuous oral immunosuppressive medications. The adjusted cumulative ISI score showed stronger biological alignment and tracked disease stage appropriately. The median adjusted cumulative ISI scores were close to 1 during the peri-diagnosis phase (except for pre-treatment encounters), dropped to around 0.5 in remission, and around 0.3 in relapse encounters across the three AAV cohorts, thereby supporting clinical plausibility. The correlation between TTV count and cumulative ISI score was slightly stronger for the adjusted score than the base (unadjusted) (r=0.37 vs 0.35; both p<0.001). Therefore, subsequent clinical analyses focused on the adjusted score. Among the GEE models, the model including the adjusted cumulative ISI score had the lowest QIC, compared with the unadjusted ISI score model and the categorical treatment indicator model, indicating better relative model fit. Conclusions: We describe, for the first time, a pragmatic ISI score to represent a patient's overall immunosuppressive treatment status at a specific time point. This provides a reusable treatment-state variable for clinical analytics and prognostic modeling and represents a first step toward biomarker-enriched digital twins of immunosuppressive state for future decision support and translational digital medicine applications.

  • Acceptability of a Freely Available App-Delivered Cessation Treatment Among Adults 60+ Years: A Longitudinal Mixed-Methods Investigation

    Background: Older adults are a high priority population for tobacco cessation. Yet, this age group commonly experiences barriers (e.g., mobility impairments, lack of transportation) to in-person evidence-based cessation treatment. App-delivered cessation programs are publicly available and an accessible modality in which to widely deliver evidence-based treatment to this population. Despite promise, there has been limited research on the acceptability and efficacy of these cessation treatments within older adult populations. Objective: To (1) examine treatment acceptability and (2) identify treatment facilitators and barriers to a publicly available app-delivered cessation program among adults 60+ years who smoke cigarettes. Methods: U.S. adults 60+ years who reported past-month cigarette use and owned a smartphone were recruited via social media. At baseline, participants reported sociodemographic characteristics, cigarette smoking patterns, quitting interest/self-efficacy, and digital literacy. Personnel instructed participants how to download a National Cancer Institute freely available cessation app, with no usage guidelines imposed. At a one-month follow-up, participants completed semi-structured interviews regarding treatment acceptability. Using a deductive-inductive thematic analysis approach, themes were identified and meaningfully organized by the Technology Acceptance Model. Subsequently, qualitative and quantitative data were integrated using the Pillar Integration Technique to create “pillars” converging mixed data. Results: Participants (N=30; age range 60-83 years) were mostly (73%) women and diverse by race, education, and income. On average, this sample was highly motivated to quit (M=8.7/10), had moderate quitting self-efficacy (M=5.4), and reported high digital proficiency (M=4.7; possible range 1-5). Participants smoked an average of 13 cigarettes per day, with the majority having moderate or high nicotine dependence. At follow-up, 67% said they would use the app in the future and almost half (46%) reported daily usage. Thematic analysis identified 10 themes overall and 5 pillars converged 2 quantitative categories with 8 qualitative themes. Individuals with low to moderate dependence described the app as useful (e.g., distraction from cravings, educational); whereas those with high dependence did not. Individuals with moderate quitting interest also described the app as useful and valued its self-guided delivery format. Those highly interested in quitting wanted more instructions for optimizing treatment. Participants with moderate to high interest in quitting described the app as easy to use and motivating; whereas those with low motivation did not. Conclusions: A publicly available app-delivered program was an acceptable cessation treatment among adults 60+ years. Acceptability was highest among individuals with moderate interest in quitting and low to moderate nicotine dependence. App-delivered treatments might be optimal for adults 60+ years who are contemplating cessation but not yet ready to engage with more intensive treatment, providing an accessible opportunity to explore quitting at one’s own pace. Studies should identify app components that may enhance acceptability among individuals with low motivation to quit and high nicotine dependence.

  • Voluntary Web Surveys Yield Higher Obesity Estimates Than Mandatory Screening in Chinese University Students

    Background: Background: Web-based surveys dominate health data collection among young adults, yet validation studies rely on mandatory participation or in-person verification, conditions absent from real-world digital surveillance. Whether voluntary web-based surveys produce systematically different estimates than mandatory objective assessment is unknown. Objective: Objective: We compared BMI from a voluntary, anonymous web-based survey with objectively measured BMI from a mandatory fitness assessment in the same university population. Methods: Methods: We paired a voluntary web-based survey (n=7,465; Wenjuanxing platform) with the mandatory Chinese National Student Physical Fitness Standards assessment (n=14,166) at a Chinese engineering university. Under full anonymity, individual matching was infeasible. We constructed six gender-by-grade strata, computed stratum-level discrepancies, and used quantile mapping and counterfactual bounding to distinguish selection from reporting effects. Bootstrap 95% CIs quantified uncertainty. Results: Results: Voluntary survey BMI exceeded mandatory assessment BMI in all six strata (+0.61 kg/m2 weighted mean). The discrepancy was driven by weight (+0.7 to +3.0 kg), not height (+0.5 to +1.0 cm). Bootstrap CIs crossed zero in the two largest strata. Self-reported obesity prevalence was 10.3% versus 8.2% measured. Treating the discrepancy as measurement error reduced obesity prevalence to 9.2%. Conclusions: Conclusions: The pattern, weight-driven, concentrated in smaller strata, indistinguishable from zero in largest strata, is consistent with heavier individuals being more likely to respond to voluntary health surveys, not with systematic reporting error. The distinction between reporting bias and selection bias determines whether the remedy is better instructions or better sampling design. Clinical Trial: no

  • Can digital storytelling enhance recovery in bipolar disorder?: A focus group study with patients and family members

    Background: Digital storytelling is an emerging approach in healthcare that blends narrative medicine with multimedia technology to share lived experiences of illness. While prior research has demonstrated benefits for creators of digital stories and for healthcare professionals who view them, less is known about how such stories impact patients and their families. Objective: This study explored how viewing a digital storytelling series about bipolar disorder influences patients and family members, particularly regarding personal recovery. Methods: We conducted a qualitative study using focus groups with patients diagnosed with bipolar I or II disorder and family members. Participants viewed a five-part digital storytelling series (Out of Darkness) and engaged in guided discussions. Data were analyzed using reflexive thematic analysis, with interpretation informed by the CHIME-D recovery framework (Connectedness, Hope, Identity, Meaning, Empowerment, and Difficulties/Trauma). Results: A total of 32 participants (17 patients and 15 family members) took part in 8 focus groups. These participants consistently described digital storytelling as emotionally impactful, relatable, and validating. Participant narratives reflected all CHIME domains. Participants described feelings of connectedness (“feeling seen and less alone”), hope, reduced stigma, strengthened identity, and greater empowerment in managing illness. The stories also prompted reflection on difficulties and trauma, which participants described as both challenging and healing. Family members reported enhanced empathy and understanding of their loved ones’ experiences. Conclusions: Digital storytelling appears to complement traditional psychoeducation by addressing emotional and experiential aspects of illness. It may support personal recovery in bipolar disorder by facilitating connection, hope, meaning, and agency while acknowledging the complexity of lived experience.