Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

The leading peer-reviewed journal for digital medicine and health and health care in the internet age. 

Latest Submissions Open for Peer Review

JMIR has been a leader in applying openness, participation, collaboration and other "2.0" ideas to scholarly publishing, and since December 2009 offers open peer review articles, allowing JMIR users to sign themselves up as peer reviewers for specific articles currently considered by the Journal (in addition to author- and editor-selected reviewers).

For a complete list of all submissions across all JMIR journals as well as partner journals, see JMIR Preprints

Note that this is a not a complete list of submissions as authors can opt-out. The list below shows recently submitted articles where submitting authors have not opted-out of open peer-review and where the editor has not made a decision yet. (Note that this feature is for reviewing specific articles - if you just want to sign up as reviewer (and wait for the editor to contact you if articles match your interests), please sign up as reviewer using your profile).

To assign yourself to an article as reviewer, you must have a user account on this site (if you don't have one, register for a free account here) and be logged in (please verify that your email address in your profile is correct).

Add yourself as a peer reviewer to any article by clicking the '+Peer-review Me!+' link under each article. Full instructions on how to complete your review will be sent to you via email shortly after. Do not sign up as peer-reviewer if you have any conflicts of interest (note that we will treat any attempts by authors to sign up as reviewer under a false identity as scientific misconduct and reserve the right to promptly reject the article and inform the host institution).

The standard turnaround time for reviews is currently 2 weeks, and the general aim is to give constructive feedback to the authors and/or to prevent publication of uninteresting or fatally flawed articles. Reviewers will be acknowledged by name if the article is published, but remain anonymous if the article is declined.

The abstracts on this page are unpublished studies - please do not cite them (yet). If you wish to cite them/wish to see them published, write your opinion in the form of a peer-review!

Tip: Include the RSS feed of the JMIR submissions on this page on your homepage, blog, or desktop RSS reader to stay informed about current submissions!

JMIR Submissions under Open Peer Review

↑ Grab this Headline Animator

If you follow us on Twitter, we will also announce new submissions under open peer-review there.

Titles/Abstracts of Articles Currently Open for Review:

  • Uncoded Clinical Features from Multilingual Electronic Health Records in Catalonia: Development and Validation Study

    Background: Unstructured free‑text narratives in electronic health records (EHRs) contain critical clinical information that is not captured by structured standard codes. In Catalonia, primary care notes follow a semi‑structured format known as MEAP (in Catalan). However, extracting structured variables using cloud‑hosted commercial large language models (LLMs) raises substantial data privacy concerns and incurs high operational costs. Objective: To develop and validate a privacy-preserving local hybrid pipeline combining a lightweight small language model (SLM) with deterministic regular expressions for extracting six uncoded urinary tract infection (UTI) clinical features (fever, nitrites, leukocytes, lumbar pain, abdominal pain, and haematuria) from primary care EHR narratives Methods: We deployed the open‑weight model (Phi4-mini, 3.8B parameters) within the institutional firewall using Ollama. Prompts were iteratively refined in collaboration with clinicians, and regular expressions were applied post‑inference. Validation was conducted across two independent arms: (1) a double‑blind clinician gold standard review of 60 real patient records (consensus‑resolved cases), and (2) an adversarial synthetic dataset of 720 notes enriched with linguistic noise, generated using GPT‑4.1 and Grok‑4.1. Point estimates and 95% confidence intervals (CIs) were computed for key performance metrics. Results: The pipeline processed 15,498 MEAP narratives from a matched cohort of 2,962 primary care patients and identified 4,663 clinical feature occurrences. Patients progressing to acute pyelonephritis (cases) presented a higher features burden than non-progressing controls (71.3% vs. 55.8%; standardized mean difference [SMD] = 0.326), particularly for fever (33.0% vs. 9.0%; SMD = 0.616) and lumbar pain (29.0% vs. 9.8%; SMD = 0.502). In real-world clinician validation, overall performance yielded 93.8% accuracy (95% CI 90.6%-96.1%), 83.6% sensitivity (95% CI 73.0%-91.2%), 96.6% specificity (95% CI 93.6%-98.4%), 87.1% positive predictive value (PPV; 95% CI 77.0%-93.9%), and 95.5% negative predictive value (NPV; 95% CI 92.3%-97.7%). In synthetic stress-testing, the pipeline demonstrated 87.2% accuracy (95% CI 84.6%-89.6%), 73.9% sensitivity (95% CI 68.9%-78.5%), near-perfect specificity (99.5%, 95% CI 98.1%-99.9%), and 99.2% PPV (95% CI 97.2%-99.9%). Conclusions: A privacy-preserving local hybrid framework combining a compact open-weight SLM with regular expression rules achieves high specificity and precision for extracting six uncoded clinical features from multilingual primary care narratives. Running entirely within institutional servers, this approach keeps patient data secure, meets privacy standards, and avoids Application Programming Interface (API) costs, making it a practical tool for EHR research, though with lower sensitivity for narratively complex descriptions.

  • Background: Internet-delivered cognitive behavioural therapy (iCBT) reduces psychological distress in patients with cardiovascular disease (CVD), but whether formats with less clinical support achieve outcomes comparable to structured therapist-directed approaches remains unclear. Objective: To evaluate the non-inferiority of self-selected versus therapist-selected content and feedback on demand versus regular scheduled therapist feedback in individualised iCBT for patients with CVD and psychological distress. Methods: This 2×2 factorial non-inferiority randomised controlled trial compared therapist-selected programs with regular scheduled feedback (TS-RF) or feedback on demand (TS-OD), and self-selected programs with regular scheduled feedback (SS-RF) or feedback on demand (SS-OD). Participants with CVD and symptoms of depression, anxiety, or stress were recruited nationwide in Sweden. Primary analyses compared therapist-selected (TS-RF and TS-OD) with self-selected programs (SS-RF and SS-OD), and regular scheduled feedback (TS-RF and SS-RF) with feedback on demand (TS-OD and SS-OD). Exploratory analyses compared each alternative format (TS-OD, SS-RF, and SS-OD) with TS-RF. Primary outcomes were depressive symptoms (Patient Health Questionnaire-9 [PHQ-9]), anxiety (Generalized Anxiety Disorder-7 [GAD-7]), and perceived stress (Perceived Stress Scale-10 [PSS-10]); cardiac anxiety (Cardiac Anxiety Questionnaire [CAQ]) was secondary. Outcomes were assessed at baseline, post-treatment (9 weeks), and 6-month follow-up. Non-inferiority used a predefined margin of d=0.30. Results: Among 237 randomised participants, all groups showed large improvements (Cohen's d >0.90) in depressive symptoms, anxiety, and stress, largely maintained at 6 months. Differences between therapist-selected and self-selected programs and between regular scheduled and on-demand feedback were generally small, but confidence intervals frequently exceeded the predefined non-inferiority margin, precluding firm conclusions. High adherence was more common in therapist-selected than self-selected programs (71.2% vs 45.4%) and with regular scheduled than on-demand feedback (64.4% vs 52.1%). On-demand feedback reduced therapist time (approximately 2 vs 6 minutes per participant per week). Exploratory comparisons of the original trial arms likewise failed to demonstrate non-inferiority of TS-OD, SS-RF, or SS-OD relative to TS-RF. Conclusions: iCBT formats with reduced clinician input may yield similar outcomes for some patients with CVD, although non-inferiority was not consistently demonstrated. Therapist-selected programs and regular scheduled feedback were associated with higher adherence. Thus, greater self-direction may increase flexibility and reduce therapist time, while regular scheduled support may remain important for engagement. Future research should identify patients who can successfully use lower-support formats. Clinical Trial: Clinical Trials.gov (registration number NCT04726722)

  • Universal guidelines for the prevention, identification and management of imposter participants in online research

    The rise of online research methods for recruitment and data collection following the COVID-19 pandemic has coincided with a growing threat to research integrity: imposter participants who falsify experiences to participate in paid studies. This viewpoint article provides a summary of the contemporary evidence on combatting imposter participants and proposes universal, evidence-based guidelines. Case reports and scoping reviews examining this threat are numerous, and some recommend guidelines for various study designs. Few studies have examined prevalence, financial implications and personal impacts. Imposter participants pose a persistent, evolving threat to online research. Responsibility cannot rest with individual researchers alone. Guidelines along with investment in institutional supports are urgently needed to safeguard data integrity and maintain public trust in research. Our recommended guidelines target each part of the research eco-system: 1) publishers, editors and checklist authors; 2) national funding and ethics bodies; 3) universities and other research institutions; and 4) researchers and study teams.

  • Digital Communication With Public Health Agencies: A Factorial Survey Study of Citizen Preferences

    Background: Public health agencies (PHAs) in Germany are increasingly exploring digital communication channels to improve citizen engagement and responsiveness. However, little is known about citizens' preferences for such digital services, especially concerning who initiates contact - whether citizens reach out to agencies or vice versa - and the conditions under which different digital channels are accepted or rejected. Privacy concerns and perceived benefits of digital services may jointly shape citizens' willingness to engage, consistent with a privacy calculus perspective. Objective: This study investigates citizens' preferences for digital information and advisory channels offered by PHAs in Germany. Specifically, we experimentally examine how service attributes (contact reason, channel, and perceived benefit) shape usage intentions for citizen-to-agency versus agency-to-citizen communication. Additionally, we explore the moderating roles of urban/rural residence, privacy attitudes, and sociodemographic characteristics. Methods: We conducted two factorial survey experiments in Saxony (Germany). The vignette experiments were embedded in a population-representative mixed-mode survey (N=2,133) of residents in a rural district (District of Bautzen (BZ), n=1,034) and an urban area (City of Dresden, n=1,099). Each experiment comprised a 3×3×3 design (27 vignettes), with usage intention rated on a 7-point scale. Multilevel regression models assessed the relative effects of vignette dimensions and respondent characteristics on usage intention, separately by region and experiment. Results: Across both experiments and regions, the digital channel was the strongest predictor of usage intention. In experiment 1 (citizen contacts the PHA), online appointment booking was strongly preferred over electronic medical history forms, while video consultation was consistently rejected. In experiment 2 (PHA contacts the citizen), a dedicated PHA app was strongly preferred over video consultations, while social media channels were strongly rejected. The reason of contact and stated benefits played comparatively minor roles. Greater informational privacy concerns significantly reduced usage intention among urban respondents in Dresden, not in the rural municipalities. Older respondents reported slightly lower usage intentions in both regions. Gender and residential context within the rural district were not associated with usage intention, indicating no systematic exclusion of specific population groups. Subjective social status (SSS) showed small regionally varying positive effects. Conclusions: Citizens reported higher usage intentions for digital PHA services such as online appointment bookings and dedicated apps, than for video consultations and social media. This pattern suggests a preference for familiar, self-directed services, requiring less personal disclosure. Limited sociodemographic differences in stated preferences suggest that equitable adoption may depend primarily on service design rather than user characteristics, while privacy concerns represent a relevant but regionally variable barrier. PHAs should prioritize low-threshold, autonomous digital services and avoid relying on social media or unsolicited general health information in saturated information environments.

  • Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

    Background: Motivational interviewing (MI) is a collaborative communication approach used to elicit autonomous motivation in health-related behavior change. Generative artificial intelligence (GenAI) provides new opportunities to deliver MI through conversational systems, but evidence on how these systems are designed, assessed, and translated into interventions remains fragmented. Objective: This scoping review aimed to characterize the current evidence on GenAI-MI chatbots across system design, safety measures, MI quality, user perceptions, and intervention outcomes. Methods: We conducted a scoping review in accordance with PRISMA-ScR. Studies published or publicly available from January 2015 to June 2, 2026 were identified through PubMed, Web of Science, Scopus, PsycINFO, DBLP, IEEE Xplore, ACM Digital Library, arXiv, and ACL Anthology. Eligible studies used generative AI to generate MI-related chatbot responses or counselor utterances. Data were extracted using a predefined extraction framework and synthesized descriptively. Results: Forty-seven reports comprising 48 studies were included. Twenty studies (41.7%) focused on system design without direct participant use, whereas 28 (58.3%) involved direct interaction with a GenAI-MI chatbot. Most systems were text based and disembodied, and 23 of 48 studies (47.9%) incorporated dynamic adaptation. Safety measures were unevenly reported, with privacy and data protection (20/48, 41.7%) and safety-oriented content generation (17/48, 35.4%) more commonly described than automated (7/48, 14.6%) or human (6/48, 12.5%) risk monitoring. Among studies involving direct participant use, 21 of 28 (75.0%) reported informed consent or user education. Thirty of 48 studies (62.5%) assessed MI quality using observer- or client-based evaluations. Existing observer and client evaluations generally suggested that GenAI-MI chatbots could produce MI-consistent interactions. User perceptions were generally favorable, particularly for empathy, usability, helpfulness, and intention to use, although measurement approaches were heterogeneous. Eighteen of 48 studies (37.5%) reported intervention outcomes, including applications in physical activity, smoking cessation, and alcohol or substance use. Most intervention studies involved a single session, and only 3 of 18 (16.7%) evaluated repeated use over 10 days to 4 weeks. Positive findings were reported more consistently for short-term motivation outcomes than for sustained behavioral or functional change. Conclusions: Current evidence suggests that GenAI-MI chatbots can deliver MI-consistent interactions that are generally perceived favorably by users, but evidence supporting sustained behavioral or functional change remains limited. Future research should strengthen runtime safety monitoring, standardize MI quality assessment, and use longer-term comparative designs with behavioral and functional outcomes to determine whether short-term motivational changes could translate into meaningful intervention effects. Clinical Trial: OSF Registries xzw4g; https://osf.io/xzw4g

  • Disclosing Environmental Costs: A Transparency Framework for Healthcare AI

    Healthcare organizations increasingly procure artificial intelligence (AI) systems whose environmental impacts extend beyond the software itself to the energy, water, chemical, and community burdens of the supporting infrastructure, yet they lack a shared starting point for evaluating this information during procurement. Existing sustainability and procurement frameworks address parts of this problem, but we did not identify one that combines healthcare-specific, pre-contract use with disclosure expectations spanning energy, water, chemical, and community-siting dimensions. Drawing on a scan of existing frameworks, we propose a draft, discussion-stage question set covering 7 dimensions: disclosure quality, governance and accountability, energy and carbon, renewable energy, water, community burden, and lifecycle and chemical accountability. For each dimension, a request-for-information instrument records a vendor's disclosure status, verification, applicability, the responsible supply-chain actor, and documented effort to obtain missing information, rather than assigning a composite score; it characterizes disclosure, not environmental performance. Because smaller or downstream vendors often depend on upstream infrastructure providers, we distinguish a vendor's capacity to disclose from its documented effort to obtain unavailable information, an equity consideration central to the framework's design. Health systems and group purchasing organizations can use this instrument now as a discussion and negotiation tool, while standards bodies and policymakers develop complementary requirements for AI vendors and data-center operators. The instrument is unvalidated: further work is needed to establish inter-rater consistency, test applicability across vendor types, and clarify how disclosure relates to independently measured environmental outcomes.

  • Accuracy of Sudden Unexpected Infant Death Prevention Content on TikTok: A Content Analysis and AAP Guideline Comparison

    • This manuscript needs more reviewers

    Background: Sudden unexpected infant death (SUID) remains a leading cause of infant mortality. The American Academy of Pediatrics (AAP) provides clear guidelines to reduce SUID risk, but to what extent these guidelines are followed in online advice remains uncertain. Objective: This study investigates how SUID prevention is portrayed on TikTok and whether the content aligns with AAP recommendations. Methods: A content analysis was conducted on 293 TikTok videos uploaded before July 2024 and selected via purposive sampling based on relevant hashtags related to SUID prevention. Videos were analyzed inductively to identify SUID recommendations, and deductively to measure frequency of occurrence of specific recommendations. Each recommendation was compared to the AAP guidelines for accuracy. Differences in adherence were assessed based on content creator type (parent/caregiver, board certified health professional, wellness consultant or undeclared), using descriptive statistics and a one-way ANOVA. Results: Approximately half (43.1%) of all TikTok SUID prevention recommendations aligned with AAP guidelines, with supine sleeping and avoiding loose bedding most frequently mentioned. A minority were partially aligned (6.2%) or directly contradicted (12.3%) AAP guidelines. Finally, 38.5% of recommendations were not mentioned in AAP guidelines. Some AAP recommendations, such as safe sleep for preterm infants or swaddling safety, were underrepresented. Adherence varied significantly by creator type, with parents being least accurate. Conclusions: While many videos support AAP-aligned practices, inconsistent and incomplete information persists. The absence of several key AAP recommendations indicates a gap in comprehensive health messaging. These findings emphasize the need for more consistent, evidence-based health communication on social media.

  • Background: Mobile phone addiction among children and adolescents is a growing global concern, with prevalence reaching 22%–32%. Most research has focused on eastern China or college students, and few studies have examined the heterogeneous risk profiles of children and adolescents in Yunnan Province, a multi-ethnic border region of southwestern China. Objective: This study aimed to characterize the prevalence of mobile phone addiction among children and adolescents in Yunnan Province, to identify distinct subgroups with different addiction profiles using k-means cluster analysis, and to describe their behavioral patterns and risk characteristics, to facilitate more targeted identification of high‑risk groups and inform future interventions and policy decisions. Methods: Validated questionnaires, including a sociodemographic survey and a mobile phone addiction screening scale, were used to collect data from children and adolescents in Yunnan Province selected by cluster random sampling. Non-parametric tests, chi-square tests, and multivariate binary logistic regression were used to examine associated factors. K-means cluster analysis was used to examine heterogeneity. The 16 prefectures were divided into a development set (12 prefectures) and an external validation set (4 prefectures). Within the development set, a prefecture-based cross-validation approach was adopted, with 11 prefectures as the training set and the remaining one as the validation fold in each iteration. The final clustering solution was applied to the external validation set. Results: A total of 44,153 children and adolescents were included. Age, sex (female vs. male), grade (Grades 5–10 vs. Grade 4), boarding status, only-child status, family size, maternal education, household income, religion, alcohol use, smoking, peer and family relationships, physical activity, and self-reported academic performance were significantly associated with mobile phone addiction. K-means clustering identified three profiles: higher-grade socially active urban students, younger day students, and ethnic minority boarding students. Cluster proportions differed by less than 1 percentage point between the development and external validation sets, and cluster-specific associations remained significant and directionally consistent in the external validation set. Conclusions: As shown by the K-means cluster analysis, three heterogeneous high-risk profiles of mobile phone addiction were identified and their robustness confirmed through external validation across geographically distinct prefectures, underscoring the need for subgroup differentiation for precision screening. These profiles may support efficient screening of at-risk children and adolescents and guide targeted psychological and behavioral assessment and intervention for high-risk groups.

  • Beyond the Therapy Hour: Smart Interpersonal Psychotherapy for Continuous Mental Health Care

    Depression is among the leading causes of disability worldwide, yet access to effective psychotherapy remains limited across healthcare systems. Interpersonal Psychotherapy (IPT) is a well-established, evidence-based treatment for depression and other mental disorders that emphasizes the bidirectional relationship between mood and interpersonal functioning. However, traditional delivery models require trained clinicians and structured therapy sessions, limiting scalability and accessibility. Digital mental health technologies offer promising opportunities to expand treatment reach, but many existing tools lack grounding in established evidence-based intervention frameworks and demonstrate limited long-term engagement. In this conceptual paper, we propose Smart Interpersonal Psychotherapy (S-IPT), a digitally augmented adaptation of IPT designed to extend traditional in-person treatment beyond the therapy session through mobile sensing, digital phenotyping, and adaptive behavioral interventions. S-IPT preserves the core theoretical structure of IPT—including the four interpersonal problem areas—while integrating advances in mobile health technologies capable of monitoring behavioral indicators of social engagement. By translating interpersonal mechanisms into digitally supported interventions, and by mapping those interventions onto each of the four IPT problem areas, S-IPT aims to reinforce therapeutic insights in real-world contexts and support relational functioning between clinical encounters. We examine how digital behavioral signals, conversational interfaces, and just-in-time adaptive interventions can support the core mechanisms of IPT while maintaining fidelity to the treatment model. Ethical considerations, including maintenance of the therapeutic alliance, and suicide risk management, are also discussed. Finally, we outline a clinical research agenda necessary to evaluate the effectiveness and implementation of digitally augmented IPT within real-world mental health systems.  

  • Background: Osteonecrosis of the femoral head (ONFH) is a major cause of hip joint dysfunction and chronic disability, substantially impairing patients’ quality of life. Patients with early-to-mid-stage ONFH urgently require safe and effective conservative treatments to delay or avoid total hip arthroplasty. As a representative therapy of traditional Chinese medicine (TCM), acupuncture is widely used in the conservative management of ONFH; However, the comparative effectiveness among different acupuncture modalities remains unclear Objective: To systematically compare the efficacy of various acupuncture therapies for early-to-mid-stage ONFH using a Bayesian network meta-analysis, thereby providing evidence-based support for optimal clinical decision-making. Methods: We systematically searched CNKI, Wanfang, VIP, SinoMed, PubMed, EMbase, Cochrane Library, and Web of Science from inception to June 20, 2026, for randomized controlled trials (RCTs) evaluating acupuncture therapies for early-to-mid-stage ONFH. Two investigators independently performed literature screening, data extraction, and methodological quality assessment using the ROB2 tool. Bayesian network meta-analyses were conducted using Stata 16.0 and R 4.3.0. For dichotomous outcomes, we calculated relative risk (RR); for continuous outcomes, mean difference (MD); all with 95% credible intervals (CrI). Interventions were ranked by the surface under the cumulative ranking curve (SUCRA), and evidence quality was assessed using the CINeMA framework. Results: A total of 34 RCTs involving 2,458 patients and 14 acupuncture modalities were included. For overall response rate, fire needling combined with pricking and bloodletting (SUCRA = 0.879) ranked first, followed by hooked-needle therapy (SUCRA = 0.823) and intensive acupuncture (SUCRA = 0.799); all three were significantly superior to simple acupuncture. For Harris hip score (HHS), acupotomy combined with Chinese herbal medicine (CHM) ranked highest (SUCRA = 0.917), followed by acupotomy plus fire needling (SUCRA = 0.864) and pricking–bloodletting (SUCRA = 0.749); acupotomy + CHM improved HHS by 11.52 points over acupotomy alone (95% CrI: 5.97–17.18). For visual analogue scale (VAS) pain score, acupotomy + CHM ranked first (SUCRA = 0.804), followed by acupotomy plus fire needling (SUCRA = 0.761) and intensive acupuncture combined with rehabilitation training (SUCRA = 0.625). The CINeMA appraisal showed that 50% (6/12) of the key comparisons were of high quality, while 41.7% (5/12) were of low quality, mainly due to imprecision and within-study bias. Conclusions:  The efficacy of different acupuncture therapies for early-to-mid-stage ONFH varies considerably. Fire needling combined with pricking–bloodletting may confer advantages in improving overall response, whereas acupotomy plus CHM appears to offer greater benefits in enhancing hip function and alleviating pain. However, the SUCRA rankings are exploratory in nature and should not be equated with confirmatory conclusions from head-to-head comparisons, as they are constrained by the quantity, quality, and network structure of included studies. Clinical decisions should integrate individual patient conditions and prioritize high-quality direct evidence. These findings warrant confirmation through further large-sample RCTs with long-term follow-up.

  • Background: Artificial intelligent agents (AIAs), capable of autonomous decision-making without human intervention, are increasingly discussed for future use in health care. However, empirical evidence on how prospective health care professionals (HCPs) evaluate ethically consequential decisions made by AIA compared to human decision-makers remains limited, particularly across different clinical settings and ethical principles. Objective: This study examined whether medical students and psychology or psychotherapy students evaluate identical medical moral dilemmas differently depending on whether decisions are made by an AIA or a human decision-maker, and whether these evaluations differ by principle of action (utilitarian vs. deontological) and medical scenario setting (somatic vs. psychiatric). Methods: In this randomized, factorial, anonymous web-based vignette study, students from Germany, Austria, and Switzerland rated 12 medical moral dilemma vignettes on acceptance and responsibility attribution (6-point Likert scales), following a full factorial 2 (decision-maker) x 2 (scenario setting) x 2 (principle of action) design. The data were analyzed in R using linear mixed models with Type III Wald tests. Results: A total of N=358 participants (M age=25.0 years, SD=7.42; 74.6% female; M semester=5.9, SD=3.5) provided 4296 scenario ratings. Decisions made by AIA were rated as less acceptable (P<.001) and were attributed less responsibility (P=.002) than identical decisions made by humans. Utilitarian decisions were more accepted and attributed more responsibility than deontological decisions (both P<.001). Acceptance was lower for psychiatric than somatic scenarios (P<.001), though this setting effect was not significant for responsibility attribution (P=.22). The acceptance gap between AIA and human decision-makers was larger in somatic than psychiatric settings (P=.02), driven by declining acceptance of human decisions in psychiatric scenarios, while AIA acceptance remained stable across settings. Responsibility attribution showed a three-way interaction between decision-maker, principle of action, and scenario setting (P<.001). In psychiatric scenarios, humans were attributed more responsibility than AIA regardless of principle of action, whereas in somatic scenarios, this gap emerged only for utilitarian decisions. Neither field of study nor other demographic variables were associated with acceptance or responsibility attribution. Conclusions: Prospective HCPs consistently rated decisions made by AIAs as less acceptable and less deserving of responsibility than identical human decisions, but this gap varied systematically by clinical setting and ethical principle rather than reflecting a uniform aversion. These findings suggest that acceptance of and responsibility attribution toward AIAs in health care are context-dependent judgments, with implications for the design and clinical integration of autonomous artificial intelligent systems.

  • Background: Telehealth is underutilized to access primary care by mid-life and older adults with chronic conditions. Emergency department over-utilization or delayed care may result from barriers to accessing primary care. We sought to address gaps in evidence around factors associated with telehealth preference over in-person care in adults with chronic conditions, and relationships between telehealth preference and delayed or forgone care or emergency department use. Objective: The primary goal of this study was to offer insights into potential disparities in telehealth preference in mid-life and older adults with chronic conditions and inform policy and clinical strategies to reduce delays in care and unnecessary emergency department utilization through telehealth expansion. Methods: We performed a secondary analysis of cross-sectional data from the 2023 California Health Interview Survey. Adult respondents with chronic health conditions who utilized telehealth in the past year were included. We used multivariable logistic regression to test associations between demographic and health characteristics, preference for telehealth, delayed or forgone care, and emergency department utilization. Results: In the sample of 5711 telehealth users, 58% preferred telehealth over in-person for future care. Those aged 60–69 years, females, and those with higher education and income had greater odds of preferring telehealth. Respondents who identified as Asian and those who were not exclusively English speakers had lower odds of preferring telehealth. Preference for telehealth was associated with lower odds of delayed or forgone care but had no significant associations with emergency department use. Conclusions: This study reveals disparities in telehealth preference among Californians with chronic conditions, shaped by sociodemographic and health-related factors. Future longitudinal studies are needed to confirm factors motivating telehealth preference, to develop telehealth programs tailored to the unique needs of aging individuals with chronic conditions and improve their timely utilization of quality health care services. Clinical Trial: Not applicable

  • Berberine for Type 2 Diabetes: A Bayesian Network Meta-Analysis of Dose-Specific Comparisons With First-Line Antidiabetic Agents

    Background: Background: Berberine is widely used off-label for type 2 diabetes mellitus (T2DM) without regulatory approval or guideline endorsement, and dose-specific comparisons with first-line antidiabetic drugs are lacking. Objective: Objective: This study aimed to compare berberine doses with placebo and standard antidiabetic agents across glycemic, metabolic, and safety outcomes. Methods: Methods: We searched PubMed, Embase, Web of Science, Cochrane CENTRAL, and ClinicalTrials.gov from inception to January 1, 2026, for randomized controlled trials (RCTs) comparing oral berberine monotherapy with placebo or any active antidiabetic drug in adults with T2DM, and performed a random-effects Bayesian network meta-analysis with SUCRA ranking and CINeMA certainty assessment. Results: Results: We included 16 RCTs with 1460 participants. Berberine 1500 mg/day produced the greatest HbA1c reduction versus placebo (mean difference −1.13%, 95% CrI −1.98 to −0.52; SUCRA 97.9%) and significantly lowered HbA1c versus metformin 1500 mg/day (−1.07%, −1.66 to −0.49) and rosiglitazone 4 mg/day (−1.12%, −2.15 to −0.08). Berberine 2000 mg/day gave the largest fasting glucose reduction versus placebo (−1.93 mmol/L, −3.59 to −0.29). Effects on lipids and body mass index were small, generally nonsignificant, and below minimal clinically important differences. Gastrointestinal events increased at higher doses, but sparse data precluded precise estimates. Conclusions: Conclusion: Berberine 1500 mg/day showed superior glycemic control versus placebo, metformin, and rosiglitazone, whereas metabolic benefits were modest and higher doses raised tolerability concerns. Confirmatory trials are needed before guideline recommendations can be made. Clinical Trial: The protocol for this systematic review and network meta-analysis was registered prospectively with PROSPERO (CRD420251088450). We conducted the review in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines for meta-analyses (Appendix 1)

  • Background: Depression and anxiety are highly prevalent among young adults, yet access to evidence-based psychological care remains limited by cost, stigma, and workforce shortages. Large language model (LLM)–based conversational agents have been proposed as scalable, low-barrier formats for delivering cognitive behavioral therapy (CBT), but rigorous randomized comparisons against established self-help interventions are scarce, particularly in non-English contexts. Objective: This study evaluated the effectiveness and user experience of a hybrid LLM-based CBT conversational agent (chatbot) for depressive and anxiety symptoms in Korean young adults, compared with CBT bibliotherapy as an active control. Methods: We conducted a 2-week, 2-arm, parallel-group randomized controlled trial. Korean young adults aged 19 to 29 years who self-reported depressive or anxiety symptoms were randomly assigned to a web-based hybrid LLM-based conversational agent (Mango) or to a CBT-based self-help book (bibliotherapy). The intervention was fully automated, with a single in-person onboarding session and no therapist involvement. The Patient Health Questionnaire-9 (PHQ-9) and Generalized Anxiety Disorder-7 (GAD-7) were self-administered online at baseline (T0), week 1 (T1), and week 2 (T2). All outcomes were analyzed using linear mixed models with an intention-to-treat approach; effect sizes were calculated as repeated-measures Cohen d (dRM) and Hedges g (gPPC2). A corresponding completer analysis served as a sensitivity analysis. Open-ended responses were analyzed thematically. The target sample (n=80) was based on feasibility rather than an a priori power calculation. Results: A total of 90 participants were randomized (experimental, n=49; control, n=41); 74 completed all assessments and adhered for at least 7 days. Baseline characteristics, including gender, age, occupation, and initial PHQ-9 or GAD-7 distribution, did not differ between groups. In the intention-to-treat analysis, both groups showed significant reductions from baseline to T2. In the experimental group, mean changes were −2.39 for PHQ-9 (95% CI −3.54 to −1.24; P<.001; dRM=−0.52) and −2.16 for GAD-7 (95% CI −3.07 to −1.26; P<.001; dRM=−0.59). Corresponding mean changes in the control group were −3.63 for PHQ-9 (P<.001; dRM=−0.81) and −3.15 for GAD-7 (P<.001; dRM=−0.83). Between-group differences in change were not statistically significant at either follow-up; at T2, the estimated differences were 1.24 for PHQ-9 (95% CI −0.44 to 2.92; P=.15; gPPC2=0.23) and 0.98 for GAD-7 (95% CI −0.34 to 2.30; P=.15; gPPC2=0.20). Thematic analysis indicated that, despite the absence of significant between-group differences, the 2 formats engaged users through different mechanisms: Mango users described contextualized emotional exploration and a perceived therapeutic alliance alongside the burden of sustained disclosure, whereas bibliotherapy users described declarative CBT learning without personalization. Conclusions: A hybrid LLM-based CBT conversational agent did not outperform CBT bibliotherapy over 2 weeks; significant within-group improvement was observed in both conditions in the completer analysis. Because the trial was underpowered and not designed for equivalence testing, these findings do not establish that the CBT conversational agent Mango is as effective as bibliotherapy. The study provides a transparent benchmark of a hybrid LLM architecture against an active comparator and qualitative evidence of distinct engagement mechanisms, informing adequately powered evaluation of future LLM-based interventions. Clinical Trial: Clinical Research Information Service (CRIS) KCT0012309; https://cris.nih.go.kr/cris/search/detailSearch.do?seq=33449

  • Background: Wearable devices support self-monitoring but rarely deliver the adaptive, individualized coaching that translates monitoring into sustained behavior change. Large language models (LLMs) could supply that coaching layer at low marginal cost, yet randomized evidence testing LLM-enhanced coaching against an active digital comparator, with a free-living behavioral outcome, is scarce. Objective: This study aimed to determine whether adding an LLM-enhanced conversational coach (HabitBot) to a blood pressure smartwatch increases free-living physical activity more than the same smartwatch combined with a standard companion app and a nurse-written exercise prescription in physically inactive adults with prehypertension, and to examine secondary physical activity, anthropometric, engagement, acceptability and theory-aligned behavioral outcomes. Methods: We conducted a 12-week, individual-level, 1:1, assessor-blind, parallel-group exploratory randomized controlled trial in Beijing, China. A total of 46 physically inactive adults aged 18-65 years with high-normal blood pressure were randomized to HabitBot or to an active comparator. HabitBot combined GPT-4o conversational coaching delivered through WeChat, progressive action and coping planning, context-linked cues, a blood pressure smartwatch and an optional peer-support group. The comparator received the same smartwatch, a standard wearable companion app and 1 nurse-written individualized exercise prescription. The primary outcome was mean daily smartphone-recorded steps at 12 weeks, analyzed by analysis of covariance adjusted for baseline. Secondary outcomes were questionnaire-assessed physical activity; body composition, blood pressure, theory-aligned behavioral determinants, engagement, acceptability and harms were exploratory. Results: All 46 randomized participants completed the 12-week assessment and were analyzed as assigned. HabitBot increased mean daily steps by 1397.6 steps/day relative to the comparator (95% CI 399.4-2395.8; P=.007). The estimate was unchanged across heteroscedasticity-robust, covariate-adjusted, change-score, randomization-inference, bootstrap and leave-one-out sensitivity analyses. Vigorous, moderate and total physical activity favored HabitBot; blood pressure did not differ. Weekly reach was sustained at 21-23 of 23 participants, although interaction volume declined. Behavioral determinants favored HabitBot after false-discovery-rate correction, but these self-reported outcomes were materially constrained by directional baseline imbalance and unblinded self-report. No serious adverse events occurred. Conclusions: LLM-enhanced coaching added to a wearable increased short-term physical activity beyond wearable access and standard app support. The trial was small, the intervention multicomponent, and the behavioral mechanism unresolved; adequately powered, component-separating trials are needed before implementation. Clinical Trial: Chinese Clinical Trial Registry ChiCTR2400085073; https://www.chictr.org.cn/showprojEN.html?proj=215626

  • Background: Background: Africa bears over 20% of the global disease burden but contributes only 3% of clinical trials. Artificial intelligence (AI) is increasingly used in trial design, recruitment, monitoring, and analysis, yet the readiness of African Research Ethics Committees (RECs) and community engagement (CE) frameworks for AI integrated research remain unknown. Objective: Objective: To systematically identify, appraise, and synthesize empirical evidence and policy guidance on (1) REC capacity to review AI enabled clinical trials in Africa, (2) documented challenges for ethical oversight, and (3) community engagement strategies that address AI specific issues. Methods: Methods: The study was a systematic review conducted using the 2020 Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) guidelines. Empirical peer-review articles on formal policy analysis on AI in clinical trials, REC function/ethics review, or community engagement in African settings published between January 2000 and April 2026 in English or French, were retrieved from seven databases: PubMed/MEDLINE, Web of Science, SCOPUS, Cochrane Library, CINAHL, African Journals Online (AJOL), and Google Scholar as well as grey literature from the WHO, African Union, EDCTP, and national REC websites. Two reviewers independently screened, extracted data, and assessed quality using the Mixed Methods Appraisal Tool (MMAT) version 2018. Synthesis was thematic with harvest plots for capacity indicators. Results: Results: Of 1,847 screened records, 33 met the inclusion criteria, with only 24% (n=8) specifically addressing the intersection of AI and Research Ethics Committees (RECs). Findings reveal critical institutional deficits: no identified African REC utilized dedicated AI protocols, specialized training, or technical advisors. Primary challenges included algorithmic opacity (82%), data extractivism (73%), and accountability gaps (67%), while AI-adapted community engagement remained rare and unevaluated (n=6). Given the predominance of descriptive data and lack of intervention evaluations, the evidence base warrants only moderate-to-low confidence. Conclusions: Conclusions: There is a critical absence of operational REC readiness for AI integrated trials in Africa. Community engagement remains conceptual rather than implemented. Immediate priorities include mandatory AI literacy training for REC members, development of African led algorithmic audit frameworks, and pilot testing of participatory CE models with embedded evaluation.

  • eHealth Literacy and All-cause Mortality among Community Dwelling Older Adults: A Prospective Cohort Study

    Background: As health information and services increasingly move to digital platforms, eHealth literacy may play an important role in health management among older adults. However, evidence on the prospective association between eHealth literacy and hard health outcomes, such as all-cause mortality, remains limited. Objective: This study aimed to examine the association between baseline eHealth literacy and all-cause mortality among community dwelling older adults in China. Methods: This prospective cohort study included 2,144 adults aged 60 years or older in Jinan, China. Baseline data were collected through face to face interviews from March to May 2021. eHealth literacy was measured using the 8-item eHealth Literacy Scale. Vital status was ascertained from March to May 2026 by linking participants to the vital registration system using their unique national identification numbers. Cox proportional hazards models were used to estimate hazard ratios (HRs) and 95% confidence intervals (CIs) for the association between eHealth literacy and all cause mortality. Sequential models adjusted for sociodemographic characteristics, health behaviors, body mass index, and chronic diseases. Kaplan Meier curves were plotted after dichotomizing eHealth literacy at the sample median. Subgroup and sensitivity analyses were also conducted. Results: Among the 2,144 participants, the mean age was 72.0 years (SD 7.0), and 1,069 (49.9%) were women. During a mean follow up of 56.8 months, 189 deaths were recorded. Participants who died had lower baseline eHealth literacy than those who survived (mean 12.28, SD 6.73 vs mean 18.07, SD 9.69; P<.001). In the fully adjusted Cox model, higher eHealth literacy was associated with a lower risk of all-cause mortality (HR=0.969, 95% CI 0.948 -0.992). Kaplan Meier analysis showed that participants with high eHealth literacy had a higher 60-month survival rate than those with low eHealth literacy (95.6% vs 84.7%). Subgroup analyses suggested that the association was more evident among participants aged 60 to 69 years (HR=0.915, 95% CI 0.856 - 0.979) and 70 to 79 years (HR 0.966, 95% CI 0.935 - 0.998), but not among those aged 80 years or older. The association differed by age group (P for interaction=.024). Sensitivity analyses yielded consistent results. Conclusions: Higher eHealth literacy was associated with lower all cause mortality among community dwelling older adults in China. These findings extend current evidence on the health relevance of eHealth literacy to long term survival and support its consideration as a modifiable factor in healthy aging. Further intervention studies are warranted to determine whether improving eHealth literacy can lead to better long term health outcomes. Clinical Trial: NA

  • Theoretical approaches in studying access to digital health services in rural contexts in high-income countries: a scoping review

    • This manuscript needs more reviewers

    Background: Digital services are increasingly used in healthcare; however, access to these services may be hindered by several factors. People residing in rural areas may face additional challenges. Rural digital healthcare access is a complex and interdisciplinary topic; to what extent this complexity is captured and theorised in extant research remains unclear. Objective: We aimed to review the theoretical approaches used in previous studies on digital healthcare access in rural areas, specifically to identify the frameworks used and the application of intersectionality. Methods: We searched MEDLINE, Embase, CINAHL Plus, and Scopus for literature related to digital health access in rural areas in high-income countries. Search results were screened, extracted, and analysed with basic content analysis. Results: We included 101 studies in the analysis. Conceptual frameworks came from research fields of technology acceptance, healthcare access, and implementation research. The most used frameworks were the Technology Acceptance Model, the Unified Theory of Acceptance and Use of Technology, Andersen’s behavioural model, Levesque’s patient-centred framework, and the Consolidated Framework for Implementation Research. Forty-two studies focused on populations with some intersecting inequities, but none mentioned intersectionality specifically. Conclusions: All frameworks have their own limitations, especially when crossing disciplines in the specific rural context. We urge researchers to acknowledge these limitations and consider interdisciplinary frameworks. There is a need to theoretically approach the issue of inequity in this topic, and we recommend using equity-based frameworks for a comprehensive and intersectional understanding.

  • Operationalizing Patient-Centric Decentralized Clinical Trials: From Fragmented Features to Integrated, Scalable Systems

    Decentralized clinical trials (DCTs) have advanced rapidly in concept and continue to grow in demand, but they remain difficult to scale in practice. The primary barrier is not simply a technology problem, but a systems problem: the absence of an integrated operating model that aligns protocol design, workforce and service delivery models, investigator oversight, regulatory and quality requirements, and participant experience all with decentralized delivery. We argue that sustainable DCT scale depends on four interdependent capabilities that must be designed into protocols and operating models from the outset: (1) orchestrated operational infrastructure designed for mobility and remote execution; (2) regulatory and quality frameworks adapted to distributed oversight; (3) community-integrated models that prioritize equitable access and participant-centered flexibility; and (4) technology and data systems that coordinate these components in real time. Drawing on real-world implementation across large-scale decentralized and hybrid clinical trials, we distill practical lessons and design requirements that distinguish reproducible models from one-off successes. We further propose that artificial intelligence is emerging as a critical enabler of this model, not as a standalone tool but as a coordinating layer that can help automate patient to trial matching, optimize operational workflows, and support continuous oversight across distributed environments. Embedding AI within the DCT architecture has the potential to transform decentralization into a scalable, learning system by augmenting human efforts and preserving the essential role of clinical professionals. By reframing DCTs as an integrated, AI-enabled operating model, here we outline a path toward trials that are more scalable, resilient, and inclusive, while maintaining scientific rigor and regulatory confidence.

  • Using inter-contact intervals to identify acute care needs in primary care patients: a secondary data analysis using electronic medical records

    • This manuscript needs more reviewers

    Background: The inter-contact interval (ICI), defined as the number of days between consecutive practice contacts, measures individual patient contact behavior but requires further validation as a useful tool for primary care research. Objective: This study applies the ICI to different electronic datasets from primary care practices to characterize individual patients’ contact patterns over time. The ICI and electronic “big data” are then used to develop and refine measures that may signal increased contact frequency and help identify acute care needs. Methods: This retrospective longitudinal observational study used electronic medical records (EMRs) from German primary care practices for a secondary data analysis, comprising (1) all patients from 155 primary care practices over 14 years, (2) all patients from one practice over 24 years, and (3) 100 patients receiving oral anticoagulation from seven practices over 7 years. ICI distributions were summarized using means, standard deviations, medians, 25th and 75th percentiles, interquartile ranges, and boxplots for a random sample. Ordinal time series analysis was used to identify increased practice contact frequency and ICI minima that best reflected a potential new need for acute care and patterns of contact behavior. Results: Among 333,610 patients with ≥2 practice contacts, 7,044,653 ICIs were identified across the three datasets. Patients had practice contacts, on average, every 38, 53, and 15 days, respectively. Nearly half of the patients had at least one contact at a 1-day interval every 3 months. ICI frequency distributions consistently showed peaks at 7 days and their multiples and at 90–93 days. Substantial variation in ICI distributions was observed within individual patients, between patients, between practices, and across datasets. Time series analyses and further refinements suggested defining a short ICI minimum as a contact with an ICI shorter than those of its direct neighbors and <8 days, potentially indicating acute care needs. Low ICI values, indicating frequent contacts, were not randomly distributed; instead, they frequently formed clusters (“islands”) of acute care within longer periods of chronic care. Patients could be divided into three groups: patients with a nearly constant consultation rhythm and no urgent practice visit; patients with 1 or more periods of intensified contact, suggesting acute care needs; and patients with frequent short-interval practice contacts. Conclusions: While a high number of contacts is a simple indicator of potentially intensified acute care needs, refined measures, such as ICI minima provide a more stringent indication of care intensity and the severity of acute care needs. Ordinal time series analysis and graphical ICI analysis can identify “islands” of frequent attendance and distinguish acute care episodes from the surrounding chronic care trajectory, demonstrating the usefulness of the ICI as a tool for longitudinal analyses of real-world primary care data. Clinical Trial: n.a.

  • Applications of Hybrid Approaches to Care, Blending In-person and Virtual Care Modalities: A Scoping Review

    • This manuscript needs more reviewers

    Background: The Inverse Care Law stipulates that healthcare resources are unevenly distributed and tend to vary inversely with actual clinical needs. Hybrid approaches, which deliberately combine in-person and virtual care, may help counteract this phenomenon in various settings, including rural healthcare, by leveraging clinician capacity across geographic regions. However, the literature on the potential of hybrid approaches to primary care has not been comprehensively mapped. Objective: This scoping review aimed to map the international evidence on applications of the hybrid approach to primary care and examine its impacts across the domains of the Quadruple Aim framework for quality improvement. Methods: We conducted a scoping review following the PRISMA extension for Scoping Reviews and the five-stage framework by Levac et al. MEDLINE, Scopus, Embase, and CINAHL were searched for studies published from 2020 onward, supplemented by a grey literature search. We included literature describing hybrid approaches to outpatient care involving primary care clinicians. Outcomes were mapped across the Quadruple Aim domains of patient experience, clinician and non-clinical staff experience, clinical impact, and economic impact. Results: Of 3,504 unique records screened, 24 peer-reviewed studies and 8 grey literature sources were included. Peer-reviewed studies were predominantly conducted in North America (16/24) and in rural or remote settings (15/24). Hybrid approaches varied considerably in design and implementation but commonly used virtual-first models and interdisciplinary teams combining off-site physicians with locally based clinicians. Clinical impact was the most frequently assessed domain (13/24) and showed the most consistently favourable findings (11/13), followed by patient experience (6/8 studies reporting favourable findings). Clinician experience showed mixed findings, while economic impact was infrequently evaluated. Conclusions: Hybrid approaches show promise as an innovative way to address gaps in healthcare access and mitigate workforce maldistribution in hard-to-serve regions, with reported benefits spanning multiple domains of the Quadruple Aim. However, evidence regarding economic impact remains limited, highlighting an important area for future research. Clinical Trial: NA

  • The Effectiveness of Digital Physiotherapeutic Scoliosis-Specific Exercise Combined with or without Bracing for Adolescent Idiopathic Scoliosis: A Systematic Review and Meta-analysis

    • This manuscript needs more reviewers

    Background: Adolescent idiopathic scoliosis (AIS) is a complex three-dimensional spinal deformity of unknown etiology, for which bracing and physiotherapeutic scoliosis-specific exercises (PSSE) are conservative treatments for individuals with AIS. However, limited access to professional personnel and specialized facilities hinders the effectiveness of PSSE training or bracing, making the integration of digital technologies a promising yet currently under-evidenced solution. Objective: This systematic review and meta-analysis aimed to evaluate the therapeutic effects of digital PSSE training combined with or without bracing on Cobb angle, the angle of trunk rotation (ATR), Scoliosis Research Society (SRS) scores, and compliance for individuals with AIS. Methods: A comprehensive literature search was conducted across PubMed, Embase, Web of Science, PEDro, the Cochrane Central Register of Controlled Trials, the China National Knowledge Infrastructure (CNKI), Wanfang, and VIP databases for studies published up to August 31, 2026. Studies investigating digital PSSE training with or without bracing were included. The risk of bias for randomized controlled trials (RCTs) was assessed using the Cochrane Risk of Bias 2.0 tool, and the certainty of evidence was evaluated using the GRADE approach. Quasi-experimental studies were assessed using ROBINS-I. A fixed-effect meta-analysis was performed using RevMan 5.4 to pool data on Cobb angle, ATR, SRS score, and compliance for individuals with AIS. Results: A total of 7 RCTs and 2 quasi-experimental studies were included in the systematic review, with 7 RCTs included in the meta-analysis. Compared with traditional PSSE training with or without bracing, digital PSSE training with or without bracing yielded a significant reduction in the Cobb angle by 4.02 degrees, showing clear superiority over traditional PSSE training with or without bracing (mean difference (MD): 4.02, 95% confidence interval (95% CI): 3.63 to 4.40, P < .001). Furthermore, digital PSSE training significantly improved SRS scores compared to traditional PSSE training (MD: 2.16, 95% CI: 1.46 to 2.86, P < .001). However, no significant difference was found in ATR improvement between digital and traditional PSSE training. Notably, digital PSSE training effectively improved patient adherence compared to traditional methods. Conclusions: This systematic review firstly demonstrates the efficacy of digital PSSE training combined with or without bracing for individuals with AIS. The findings indicate that digital PSSE training with or without bracing for individuals with AIS can effectively improve Cobb angles, SRS scores, and compliance compared to the traditional conservative training model. Future research should focus on optimizing digital healthcare modalities combined with PSSE or bracing, and conducting multi-center, large-sample clinical trials to validate short-, medium-, and long-term therapeutic outcomes. Clinical Trial: PROSPERO CRD420261497979;https://www.crd.york.ac.uk/PROSPERO/view/CRD420261497979

  • What a Multimodal Result Establishes: Viewpoint on Three Reporting Gaps That Current Clinical-AI Guidance Does Not Cover

    Multimodal clinical decision support combines imaging, clinical text, structured records, and molecular data. It is usually judged by a single comparison: whether the fused model outperforms a model built on one modality. That comparison establishes less than it appears to. A higher score can reflect one dominant modality, an availability pattern rather than clinical signal, or a dataset shortcut, and it is measured under an input condition that deployment does not reproduce. In this viewpoint, we ask which of the reporting requirements a multimodal claim depends on are already met by current clinical-AI guidance, and which are not. We mapped 10 audit dimensions against the complete published item sets of the 3 most recent deployment-facing documents: FUTURE-AI, PROBAST+AI, and STARD-AI. Six dimensions are already covered, and work on them adapts existing guidance. One, robustness to a missing modality, sits beside an existing item that addresses a different construct. Three have no corresponding item in any of the 3 documents, and a single cause runs through them: none of these documents treats the individual modality as a unit of analysis. We therefore highlight 3 reporting gaps. First, modality ablation, because showing which inputs a model used differs from showing that it still works when one is unavailable. Second, claim provenance, because existing traceability items govern the accountability log and the origin of the dataset, leaving open whether a generated statement can be traced to its supporting evidence. Third, contradiction handling, because no item addresses 2 modalities that are individually clear and mutually inconsistent. We illustrate each gap against published systems, name the nearest existing item so that the distinction is visible, and set out the reporting that would close it. The missing items belong inside the guidance that already exists.

  • Personalized Telephone Coaching for Depression Prevention in Farmers: A Health Economic Evaluation Alongside a Randomized Controlled Trial

    Background: Major depressive disorder (MDD) is a leading contributor to the global burden of disease, associated with substantial impairment in daily functioning and reduced quality of life for the individual and considerable societal costs. Farmers constitute a particularly vulnerable population, facing an elevated depression risk due to occupational, financial, and social stressors, and simultaneously encountering considerable barriers to mental health care. Personalized telephone coaching offers a promising solution for indicative prevention in this group. However, despite the evidence on its clinical effectiveness, the cost-effectiveness and cost-utility of telephone coaching for preventing MDD remain insufficiently examined. Objective: This economic evaluation aimed to evaluate the cost-effectiveness and cost-utility of a personalized telephone coaching (IG) provided by psychologists compared to treatment-as-usual plus informational material (TAU+) in farmers from a societal and social insurance perspective within a time horizon of 18 months. Methods: An economic evaluation was performed alongside a pragmatic randomized controlled trial (RCT). Health service use, patient and family expenses, and productivity costs were assessed with the adapted German version of the self-report questionnaire TiC-P (Trimbos Institute and Institute of Medical Technology Questionnaire for Costs Associated with Psychiatric Illness). Outcomes were measured in terms of depression symptom-free status (Quick Inventory of Depressive Symptomology – Self-Report (QIDS-SR16)<6) and quality-adjusted life years (QALYs; Assessment of Quality of Life (AQoL-8D)). All assessments were conducted web-based. A cost-effectiveness (CEA) and cost-utility analysis (CUA) were performed from the societal and social insurance perspective. Analyses followed the intention-to-treat principle with multiple imputation. Uncertainty around incremental cost-effectiveness ratio (ICER) estimates was characterized using bootstrapped seemingly unrelated regression equations (2,500 simulations). Sensitivity analyses assessed the robustness of the findings. Results: In total, 314 participants were included in the RCT. From a societal perspective, the IG produced a greater number of participants achieving depression symptom-free status (ΔE=0.11, 95%CI –0.01 to 0.23) compared to TAU+ at higher costs (ΔC=€663, 95%CI 2820 to 4052; €1=US $1.12 as of 2019). The probability of the IG being cost-effective was 37% at a willingness-to-pay (WTP) threshold of €0 per depression symptom-free status and QALY gained, rising to 63% when applying a societal WTP of €20,000 per QALY gained (38% from a social insurance perspective). Analyses conducted from the social insurance perspective, as well as sensitivity analyses, indicated the dominance of TAU+ over the IG, with intervention costs as the primary cost driver. Conclusions: Findings suggest that MDD prevention via personalized telephone coaching yield larger effects, albeit accompanied by higher costs from both a societal and social insurance perspective within the observed time horizon. These results inform evidence-based decision-making on cost-effective MDD prevention in occupational groups, and highlight the need for comparative trials with alternative approaches. Clinical Trial: German Clinical Trial Registration: DRKS00015655. Registered on October 2, 2018.

  • Background: As healthcare shifts toward disease prevention and whole-life-cycle health management, artificial intelligence (AI) offers key technological enablers for proactive health management. However, medical education lacks systematic curriculum models that integrate real-world proactive health challenges with AI applications. This study aims to develop and iteratively refine an interdisciplinary curriculum framework, titled "AI-Empowered Proactive Health," based on a "Teacher–Student–AI" tripartite educational ecosystem. Objective: To address this gap, this study aims to develop and iteratively refine an interdisciplinary curriculum framework, titled “AI-Empowered Proactive Health,” based on a “Teacher–Student–AI” tripartite educational ecosystem. Methods: A sequential, three-phase iterative curriculum development method was conducted from March to December 2025. In Phase 1, an initial curriculum framework was developed via expert panel meetings using the Modified Nominal Group Technique (NGT). In Phase 2, semi-structured interviews were conducted with senior nursing undergraduates for formative evaluation. In Phase 3, two rounds of expert meetings were held to review feedback item-by-item, leading to the refinement of the 25-credit-hour curriculum content, real-world project workflows, multidimensional assessment schemes, and the online digital learning platform architecture. Results: The finalized curriculum is grounded in a "1-2-3-4-5" framework: guided by active health competencies (1 Leading Force), driven by healthcare transformation and AI advancement (2 Driving Forces), executed through instructor–student–AI collaboration (3 Roles), spanning four project phases (Problem Definition, Evidence Integration, Knowledge Graph & Agent Development, and Outcome Presentation), and integrating across five interdisciplinary domains. Student feedback resulted in reducing projects from 11 to 6 and extending contact hours from 18 to 25. The digital platform incorporates ten integrated modules (e.g., Evidence-Based Practice Workbench, Knowledge Graph Laboratory, Intelligent Agent Laboratory) to support seamless workflow execution. Assessment spans process tracking (40%), project outcomes (45%), and critical AI literacy (15%). Conclusions: Grounded in a tripartite educational ecosystem, the interdisciplinary "AI-Empowered Proactive Health" course provides a structured, practice-oriented model that bridges medical knowledge, evidence-based methodologies, and AI technology. This curriculum framework offers a valuable reference for cultivating proactive health professionals and reforming medical education in the digital era. Clinical Trial: the Ethics Committee of Beijing University of Chinese Medicine (Ethics Approval No.: 2025BZYLL0103)

  • Human-Robot Collaboration in Long-Term Care: A Scoping Review

    Background: As care robots (CRs) are increasingly integrated into long-term care (LTC), human-robot collaboration (HRC) is becoming relevant to care delivery. However, how care staff and informal caregivers collaborate with CRs, how caregiving work changes, and the associated outcomes and challenges remain unclear. Objective: This review aimed to synthesize evidence on HRC in LTC, focusing on task allocation, changes in caregiving work, and associated outcomes and challenges. Methods: This scoping review followed Joanna Briggs Institute (JBI) methodology and PRISMA-ScR guidelines. English and Chinese literature published from database inception (2005) to May 2026 was searched in CINAHL, Embase, Scopus, PubMed, Web of Science, China National Knowledge Infrastructure, Wanfang Data, and Chinese Biomedical Literature Database, supplemented by grey literature. Two reviewers independently screened studies, and data were synthesized narratively. Results: Twenty-five studies were included, of which 21 were qualitative. Most were conducted in institutional LTC settings, with Japan contributing the largest number of studies (8/25, 32%), followed by Australia (4/25, 16%) and Canada (4/25, 16%). Care robots supported routine care (e.g., reminders, logistics, and safety monitoring), psychosocial care (companionship, activity facilitation, and telepresence), and physical assistance. Two distinct human–robot collaboration models were identified: robot-intensive and human-intensive collaboration. Across both models, care robots assumed structured, repetitive, and physically demanding tasks, while care staff remained responsible for supervision, coordination, and individualized care, indicating a redistribution rather than replacement of care work. HRC was associated with reduced workload for care staff, improved well-being among care recipients, and strengthened relationships between care staff, care recipients, and family members. Key implementation challenges included technical limitations, training needs, limited adaptability across care contexts, unclear role boundaries, ethical and safety concerns, and resource constraints. Conclusions: Rather than replacing human care, care robots may reshape how care is delivered by redistributing tasks between robots and care staff. Further research should focus on designing and evaluating sustainable HRC models that maximize efficiency while preserving person-centered care. Clinical Trial: This scoping review has been registered in the Open Science Framework. The link is https://osf.io/4xdcr/overview.

  • Effects of AI-HEALS Intervention on Medication Adherence in Young and Middle-Aged Patients with Hypertension: A Randomized Controlled Trial

    Background: Young and middle-aged patients with hypertension often exhibit low rates of disease awareness, treatment, and blood pressure control, with poor medication adherence further increasing their risk of cardiovascular events. Digital intelligent interventions grounded in behavioral change theories may offer a promising approach to improving medication-taking behaviors and enhancing treatment adherence. Objective: This study aimed to develop an Artificial Intelligence-based Healthcare, Education, Assistance, and Life-management System (AI-HEALS) intervention guided by the Behavior Change Wheel (BCW) theory and to evaluate its effects on medication adherence among young and middle-aged patients with hypertension. Methods: A single-center, randomized controlled trial was conducted. A total of 82 eligible patients aged 18–59 years were randomly assigned in a 1:1 ratio to the intervention group (AI-HEALS intervention) and control group (routine management). Medication adherence, medication literacy, medication adherence self-efficacy, medication beliefs, social support, health-related quality of life, and blood pressure were assessed at baseline and at week 12. All analyses were conducted according to the intention-to-treat principle, with multiple imputation used to address missing data. Statistical analyses were performed using SPSS version 26.0. Results: Baseline characteristics were comparable between the two groups (all P > 0.05). At week 12, the intervention group demonstrated significantly higher total medication adherence scores (t = −2.376, P = 0.020), as well as higher scores in the medication cooperation (t = −2.945, P = 0.004) and intention (t = −2.189, P = 0.032) dimensions, compared with the control group. The intervention group also had significantly higher critical medication literacy (t = −2.224, P = 0.029) and medication adherence self-efficacy (t = −2.165, P = 0.033). Systolic blood pressure decreased significantly within the intervention group (t = 2.865, P = 0.007), whereas no significant between-group differences were observed in diastolic blood pressure, medication beliefs, social support, or health-related quality of life (all P > 0.05). In the control group, only the regular medication dimension and family support showed modest improvements. Conclusions: The BCW theory-guided AI-HEALS intervention effectively improved medication adherence, critical medication literacy, and medication adherence self-efficacy among young and middle-aged patients with hypertension, and was associated with a significant within-group reduction in systolic blood pressure. These findings suggest that AI-HEALS may serve as a feasible digital self-management tool for improving medication-related behaviors in this population.

  • Understanding Factors affecting User-Initiated Engagement in a Digital Mental Health Chat Platform: A Mixed Methods Analysis

    Background: Chat-based online platforms have been shown to be effective intervention tools for mental health support. However, user attrition may limit their capacity to provide meaningful support and hinder the validation of their clinical benefits. Factors contributing to the continued use or non-use of chat-based platforms must be investigated to address barriers to engagement. Objective: This study investigates user engagement with a chat-based mental health platform. Specifically, it aims to identify the factors that contribute to the initiation and subsequent initiation of sessions with mental health professionals (MHPs) through the chat platform. The findings may inform the development, refinement and implementation of chat-based mental health interventions with improved user retention. Methods: To investigate user engagement, the study employed a prospective longitudinal design. A total of 119 students were enrolled in sequential batches, with each batch observed for a three-month period. During this period, data on chat-session activity and interactions with MHPs, who are licensed clinical psychologists, were automatically recorded and collected through the SAGIP mobile app. Each MHP handled no more than 10 students at any given time. Data pertaining to participants’ well-being, mental health, demographic information and experiences with the chat intervention were collected through surveys and interviews. Of the 119 participants, 32 were interviewed, and four of them initiated a subsequent session. A binomial mixed-effects generalized linear model was used to identify factors associated with the initiation and subsequent initiation of sessions with MHPs through the online mental health chat platform. Descriptive analyses were conducted to further examine these factors and their relationships with user engagement. In addition, thematic analysis was applied to the qualitative text data to explain other factors that may influence user engagement. Results: Based on the binomial mixed-effects generalized linear model, the number of user messages sent within the current session emerged as a key significant factor associated with user initiation in that same session. Meanwhile, the number of fast responses (i.e., responses provided within five minutes) users received during their current session was a significant factor associated with subsequent initiation. The thematic analysis further revealed three critical factors influencing user engagement: the availability of mental health professionals, therapeutic alliance, and usability of the platform. Conclusions: This study combined a binomial mixed-effects generalized linear model with thematic analysis that provided a comprehensive understanding of user-initiated engagement with a mental health chat-based platform. The findings suggest that users are more likely to initiate sessions when they have concerns, experiences, or needs they wish to share, and are more likely to continue using the platform when they receive prompt and useful responses. In addition, platform usability emerged as an important factor influencing engagement. Clinical Trial: WMSU REOC Code: 2025-EF-0379

  • AnxEMU-VR: Randomized Controlled Trial Evaluating the Impact of Virtual Reality Exposure Therapy on Epilepsy-Specific Interictal Anxiety in People with Epilepsy

    • This manuscript needs more reviewers

    Background: Anxiety is the most common psychiatric comorbidity among people with epilepsy, yet epilepsy-specific (ES) interictal anxiety remains underrecognized and undertreated, highlighting the need for accessible non-pharmacological approaches. Virtual reality exposure therapy (VR-ET) can deliver controlled exposure to seizure-related fears that are difficult to reproduce in vivo, but its effects and delivery requirements need evaluation in the context of seizure unpredictability and heterogeneous seizure experiences. The epilepsy monitoring unit (EMU) allows repeated engagement under medi Objective: This mixed-methods randomized controlled trial evaluated the preliminary effectiveness of VR-ET for ES-anxiety compared with neutral VR, together with feasibility, acceptability, usability, safety, tolerability, and scenario fidelity among people with epilepsy admitted to an EMU. Methods: Fourteen people with self-reported epilepsy- or seizure-related anxiety were randomized to VR-ET (n=7) or neutral VR (n=7). Participants completed self-administered 5-minute 360-degree video sessions twice daily for up to 10 days or until discharge. VR-ET participants progressed through personalized hierarchies of graded seizure-related exposure scenes selected at baseline using a diagnostic interview; control participants viewed nature-based scenes in a standardized order. Quantitative outcomes included anxiety, fear and situational avoidance, quality of life, simulator sickness, usability, and presence, and were summarized descriptively. Post-intervention and 1-month interviews were analyzed using framework analysis with inductive coding within predefined domains. Findings were integrated through joint interpretation and cross-case comparison. Results: Quantitative outcomes were heterogeneous and showed no clear between-group pattern. Adherence was high, with all VR-ET participants and 5 of 7 (71.4%) neutral VR participants meeting the predefined threshold. Five of 7 (71.4%) VR-ET participants described the intervention as helpful, and 4 of 7 (57.1%) reported experiences consistent with exposure-related learning, including increased insight, self-efficacy, preparedness, cognitive reappraisal, and habituation. Fear and anxiety changes were variable, and incomplete 1-month follow-up limited interpretation of longer-term outcomes. Scenarios were generally realistic and representative of epilepsy-related fears, but emotional engagement depended on personal relevance. Seizure simulations were identified as most impactful, and participants recommended greater intensity, variety, interactivity, and treatment duration. VR-ET participants more often expressed interest in continued use outside the EMU, whereas neutral VR participants primarily described short-term calming and distraction during admission with some reporting proactive use during periods of heightened anxiety. Across groups, usability was favorable, simulator sickness was minimal, and no serious adverse events were attributable to VR. Engagement was affected by EMU-related fatigue, seizure activity, environmental interruptions, and technical problems. Conclusions: The VR interventions were feasible, acceptable, and safe in this sample. VR-ET showed no clear quantitative advantage over neutral VR, but qualitative findings were consistent with exposure-related learning. Neutral VR may provide low-intensity anxiety support during EMU admission. Future research should evaluate more personalized, varied, and interactive VR-ET and examine integration with therapist-supported care. Clinical Trial: ClinicalTrials.gov NCT06028945

  • Background: Assistive technology is increasingly recognized for its potential to support older adults amid increasing pressure on the care workforce, yet successful implementation in practice remains difficult. Care organizations play an important role in the selection, implementation and overall coordination surrounding the use of assistive technology in care settings, and technology suppliers can support them through the development of fitting technology and during implementation. Collaboration between them may facilitate uptake, but existing research provides only a fragmented understanding of how care organizations collaborate with suppliers. Objective: This study aimed to explore how innovation and care managers perceive and evaluate assistive technology suppliers, how these perceptions are formed, and how suppliers and care organizations collaborate in the development and implementation of assistive technology. Methods: The study employed a qualitative design using semi‑structured individual interviews with managers from organizations providing care for older adults, working in the same geographic region in the Netherlands. A semi-structured interview guide was used to guide the interviews. Interviews were audio-recorded, transcribed and analyzed using reflexive thematic analysis. Results: A total of 14 participants across 11 care organizations participated in the study. Five interconnected themes were derived pertaining to the collaboration between care organizations and assistive technology suppliers, including (1) collaboration with suppliers is shaped by individuals, (2) articulating product-related criteria for supplier evaluation, (3) seeking supplier dependability before and beyond collaboration, (4) balancing organizational oversight, autonomy, and supplier responsivity, and (5) technology development through ongoing supplier-care organization interaction. Conclusions: Findings show that care and innovation managers evaluate assistive technology suppliers not only through explicit product- and supplier-related criteria, but also through more implicit cues derived from interpersonal interactions and supplier reputation within the sector. Together, these considerations contribute to a broader assessment of suppliers as trustworthy and legitimate collaboration partners. Care organizations must navigate trade-offs within these collaborations, which may benefit from explicit consideration. More structured approaches may also help, albeit proportionate to the technology and organizational context. Sustainable supplier relationships may likewise require realistic commitments and concessions from both parties. Finally, existing supplier relationships can provide a foundation for ongoing technology development. Future research could explore forms of co-design that enable both care organizations and suppliers to contribute in ways that are feasible, meaningful, and mutually beneficial.

  • Ethical Tensions in AI-Mediated Clinical Communication Among Nursing Students: A Longitudinal Mixed Methods Study

    Background: Generative artificial intelligence (AI) is increasingly shaping how patients understand health information before clinical encounters, creating challenges for patient-professional communication involving professional authority, responsibility, uncertainty, and trust. However, little is known about how nursing students interpret these issues as they move from conceptual preparation into clinical practice. Objective: This study aimed to examine how nursing students’ interpretations of ethical issues in AI-mediated clinical communication were maintained, challenged, or reconsidered across classroom engagement, clinical placement, and subsequent reflection. Methods: We conducted a prospective longitudinal mixed methods study involving 32 third-year undergraduate nursing students, including 16 international and 16 local Chinese students enrolled in the same Bachelor of Nursing program. Questionnaire data were integrated with post-session written responses, clinical reflective logs, and post-placement group reflections across sequential classroom, clinical, and reflective contexts. Qualitative evidence constituted the primary interpretive strand, while questionnaire findings provided complementary contextual evidence. Participant-level evidence was linked longitudinally and analyzed using reflexive thematic analysis, within-case interpretation, and cross-case comparison. Results: Among the 32 participants, clinical experience complicated rather than simply extended interpretations expressed after classroom engagement. Participants encountered tensions involving the negotiation of professional and AI-informed interpretations, professional responsibilities beyond factual correction, and the communication of uncertainty while maintaining patient trust. Participant-linked longitudinal analysis identified recurring patterns of broad reconsideration across ethical issues, selective reconsideration in particular areas, and persistent or unresolved tensions. Clinical experience and subsequent reflection therefore sometimes made competing ethical considerations more visible without producing greater certainty or uniform change. Conclusions: AI-mediated clinical communication presents challenges extending beyond the accuracy of AI-generated health information to the ways in which AI-informed patient understandings interact with professional authority, responsibility, uncertainty, and trust. Preparing nurses for AI-mediated clinical communication may therefore require attention not only to the evaluation of AI-generated information but also to the relational and ethical tensions that arise when such information enters clinical encounters.

  • Background: Mobile health (mHealth) interventions are increasingly used to support personalized self-management in type 2 diabetes (T2D), particularly for promoting physical activity. Mobile apps enable self-monitoring of various indicators (eg, blood glucose, physical activity), supporting the delivery of personalized feedback and guidance. Such guidance may be provided either by health professionals (human-delivered) or automatically generated by algorithms (automated). While both approaches are promising, evidence comparing human with automated guidance remains limited. Addressing this gap is important for understanding how different approaches to personalized guidance may influence intervention effectiveness as digital health technologies are increasingly integrated into T2D care. Objective: To compare the effectiveness of human-delivered and automated personalized guidance in mHealth-based physical activity interventions for adults with T2D, and to examine physical activity outcomes, intervention engagement, and intervention characteristics. Methods: We conducted a systematic review in accordance with PRISMA guidelines. We searched four databases (MEDLINE, PsycINFO, CENTRAL, and Scopus) and two trial registries (ClinicalTrials.gov and WHO ICTRP) up to April 2026. We included randomized controlled trials (RCTs) evaluating smartphone-accessible physical activity interventions in adults with T2D that used either human-delivered or automated personalized guidance. We pooled HbA1c, the primary outcome, using random-effects meta-analysis with a generic inverse-variance method. Physical activity outcomes and intervention engagement were narratively synthesized because meta-analysis was not feasible. Results: Twenty-two RCTs (Human-delivered guidance: 10, n = 1,369; Automated guidance: 12, n = 3,135) were included. Human-delivered guidance showed a small pooled reduction in HbA1c compared with control conditions (MD -0.14%, 95% CI -0.29 to -0.00, I2=23%), which was attenuated and no longer statistically significant after excluding studies at high risk of bias. Automated guidance showed a larger pooled reduction (MD -0.31%, 95% CI -0.48 to -0.13, I2=79%) that remained statistically significant in sensitivity analysis. The certainty of evidence was low for both human-delivered and automated interventions. Evidence for physical activity outcomes remained inconclusive, with mixed findings across studies. Conclusions: Human-delivered guidance showed a smaller and less robust HbA1c effect, whereas automated guidance showed modest and more consistent improvements. The available evidence did not indicated that human-delivered guidance consistently conferred greater glycemic benefit than automated guidance. Guidance source should therefore be considered alongside delivery strategy and broader intervention content. Clinical Trial: PROSPERO: CRD420251039711

  • Multi-Perspective Chain-of-Thought Reasoning for AI-Assisted Shared Decision Support in Cannabis-Related Care: A Simulation-Based Expert Evaluation

    Background: Large language models (LLMs) may enhance AI-assisted shared decision support by translating behavioral and contextual information into complementary patient- and clinician-facing messages. However, it remains unclear whether the perceived quality of such communication differs across prompting strategies in cannabis-related care. Objective: This study compared direct generation, standard Chain-of-Thought (CoT), and Multi-Perspective CoT (Multi-CoT) across multiple LLM backbones with respect to expert selections for patient helpfulness, clinical expertise, empathy, and overall preference in simulated cannabis-related shared decision support scenarios. Methods: We developed an AI-assisted shared decision support prototype that integrates behavioral context, community-experience retrieval, emotion-aware prompting, explainable artificial intelligence (XAI) summaries, and role-specific response generation. Four blinded evaluators with diverse health care–related backgrounds each assessed 50 simulated cases, including 20 shared and 30 evaluator-specific cases, yielding 200 case-level assessments. Each evaluator was assigned a distinct pair of LLMs. For each case and evaluation criterion, evaluators selected 1 of 6 response pairs, each comprising a patient-facing message and a clinician-facing summary generated by the 2 assigned LLMs under the 3 prompting strategies. Results were summarized descriptively using counts and percentages. Results: Across 200 case-level assessments, Multi-CoT received the most selections for overall preference (122/200, 61.0%) and patient helpfulness (100/200, 50.0%). Multi-CoT and standard CoT received the same number of selections for clinical expertise (83/200, 41.5% each), whereas standard CoT received the most selections for empathy (95/200, 47.5%). Selection patterns varied substantially across evaluator–model pairings; Multi-CoT’s overall-preference selection rate ranged from 12% to 92% across the 4 evaluations. Conclusions: In this simulation-based expert evaluation, Multi-CoT was favored in aggregate for patient helpfulness and overall preference but not for empathy, and its selection patterns varied markedly across evaluator–model pairings. These findings support Multi-CoT as a promising, context-dependent prompting approach for the response-generation layer of AI-assisted shared decision support. They do not establish clinical effectiveness, improved shared decision-making, or readiness for clinical use. Further work should include multi-stakeholder evaluation, real patient–clinician interactions, and formal assessments of factual accuracy and safety.

  • Decision Aids Supporting Participation in People With Mild Cognitive Impairment or Mild-to-Moderate Dementia: A Scoping Review

    Background: People with mild cognitive impairment or mild-to-moderate dementia may retain the ability to participate in preference-sensitive decisions, although cognitive decline can make complex healthcare decision-making increasingly challenging. Decision aids may support meaningful participation, but their characteristics, development, evaluation, and implementation in this population have not been systematically synthesized. Objective: This scoping review aimed to provide an overview of decision aids for people with mild cognitive impairment or mild-to-moderate dementia, including their characteristics, applications, evaluation, and existing challenges. Methods: Our scoping review was conducted following the Joanna Briggs Institute recommendations for scoping reviews and Arksey and O’Malley’s framework. The process followed PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines and PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Literature Search Extension) checklist. A comprehensive literature search was conducted across 6 databases: PubMed, Web of Science, Cochrane Library, IEEE Xplore, CINAHL, and EMBASE. Two reviewers independently screened titles, abstracts, and full texts via Rayyan, with disagreements resolved by a third reviewer. Results: Eleven studies were included, of which 10 were published from 2019 onward. Decision-support tools addressed a broad range of decisions across the care trajectory, including driving, diagnosis and prognosis, treatment and medication management, risk reduction, and future care planning. Visual formats predominated (7/11), and more than half of the tools were designed for fully self-guided use (6/11). Most tools remained at an early stage of development (10/11), with evaluations focusing primarily on usability, feasibility, and acceptability; only three studies assessed efficacy or effectiveness. Key challenges included limited integration into clinical workflows and EHR systems, digital accessibility, organizational readiness, and limited evidence across diverse sociocultural and resource settings. Conclusions: Decision-support tools for people with cognitive impairment are expanding across a broader range of decision contexts, reflecting growing attention to person-centred and supported participation across the care trajectory. However, most tools remain at an early stage of evaluation, with limited evidence on effectiveness, decision quality, long-term outcomes, and implementation in routine clinical practice. Future research should prioritize rigorous real-world evaluation, standardized decision-quality measures, and integration into clinical workflows across diverse populations and care settings.

  • Validation of a Digital Lifestyle Assessment Tool Against Established Instruments and Its Association with Cognitive Performance

    Background: The Smart Tracker is a brief online questionnaire measuring six lifestyle domains relevant to brain health, including Diet, Exercise, Cognitive Engagement, Social Engagement, Emotional Well-being (sleep and stress/depression), and Physical Well-being. Objective: Our primary objective evaluated criterion validity of the Smart Tracker tool against established instruments measuring similar constructs. We additionally assessed the Smart Tracker’s association with objective cognitive performance in a sample subset who completed the Cogniciti Brain Health Assessment (BHA). Methods: The validation sample included 570 adults (mean age 71.3, SD 9.5 years, age range 22-97; 68.9% [393/570] female) who completed the Smart Tracker alongside five validated instruments: 1) the Eating Pattern Self-Assessment (EPSA; dietary quality), 2) the Short Form International Physical Activity Questionnaire (IPAQ; exercise), 3) the Pittsburgh Sleep Quality Index (PSQI; sleep quality), 4) the Depression Anxiety Stress Scale (DASS-21 short form; psychological wellbeing), and 5) the Cognitive Reserve Index Questionnaire (CRIq; cognitive and social engagement). A nested sub-sample (n=200) additionally completed the BHA. Criterion validity was assessed using Spearman rank correlation coefficients (ρ values) with Benjamini-Hochberg false discovery rate correction to assess the strength and direction of association between each Smart Tracker domain and the corresponding validated instruments. We used a linear regression to assess the association between the Smart Tracker and BHA. Results: The Smart Tracker questionnaire domains showed significant moderate to high correlations with the comparator validated questionnaires assessing similar constructs. Spearman correlations were significantly positive for domains related to Diet (ρ=0.62), Exercise (ρ=0.53), Cognitive and Social Engagement (ρ=0.33), and negative in the expected direction for Emotional Well-Being related to sleep (ρ=-0.65), Stress and Depression (ρ=-0.70; all P<.001). The Smart Tracker Overall score significantly predicted BHA age-normed performance in the cognitive sub-sample (Y = 36.30·X + 47.80; R²=.043; F(1,198)=8.9, P=.003) where each 10-point increase in the Smart Tracker overall score corresponded to approximately 3.6 BHA percentile points. Meaningful ceiling effects were observed for the Smart Tracker Sleep (24.9%) and Stress/Depression (31.8%) sub-scales. Conclusions: The Smart Tracker demonstrates acceptable criterion validity across each of the lifestyle domains it measures and shows a modest but significant association with objective age-adjusted cognitive performance.

  • Managing Uncertainty in AI-Enabled Occupational Injury Claims Review: An Observed-to-Expected Interpretation Framework

    Background: Health and insurance administrators increasingly use statistical models and artificial intelligence methods to identify injury or work categories for closer review. An observed-to-expected (O/E) ratio compares total recorded sickness absence days in a category with the total expected by a prediction model. A ratio greater than 1 indicates more recorded than expected sickness absence. However, it does not explain why the difference occurred or what should be reviewed next. Interpretation may depend on the prediction model, the work information defining the expected comparison, and whether the difference is concentrated among claims with very long durations. Objective: This study aimed to identify occupational injury categories with more recorded sickness absence than expected, compare their O/E results across prediction models and definitions of the expected comparison, and determine whether positive individual claim differences were concentrated among claims with very long durations. Methods: We conducted a retrospective observational study of 86,305 closed occupational injury insurance claims in Hong Kong from 2005 to 2024. Repeated cross fitting with Tweedie generalized linear models estimated expected sickness absence for each claim. Five injury and work category groupings were examined. For each grouping, its defining variable was omitted, so expected duration was based on other recorded characteristics. Categories required an O/E ratio greater than 1.10, a mean recorded minus expected difference of at least 5 days per claim, and at least 500 claims. The selected categories were assessed by replacing the Tweedie model with CatBoost after overall scale alignment, omitting available work information from the expected duration model, and estimating the share of summed positive individual claim differences contributed by claims above the empirical 95th percentile of recorded duration. The first two assessments used a predefined criterion combining relative and absolute changes. Results: Overall recorded and expected sickness absence totals were closely aligned, but 14 categories met all selection criteria. Although CatBoost had lower overall prediction error, only 2 of 14 category comparisons met the predefined change criterion after replacing the Tweedie model with CatBoost. Omitting available work information met the criterion in 6 comparisons. In 10 of 14 selected categories, claims in the upper 5% of recorded durations contributed a greater share of summed positive individual claim differences than the corresponding share in the full cohort. Conclusions: An elevated O/E ratio can identify categories for closer review but cannot explain the observed difference or determine an organizational response. Examining the same selected categories through these three questions separated category selection from subsequent interpretation. These questions provide a transparent way to examine additional evidence before comparisons generated by models are interpreted in organizational review.

  • Early versus late internet-based cognitive behavioral therapy after MINOCA or takotsubo syndrome: one-year follow-up of a randomized trial

    • This manuscript needs more reviewers

    Background: Cardiovascular (CV) events such as a myocardial infarction with non-obstructive coronary arteries (MINOCA) or takotsubo syndrome (TS) are associated with heightened levels of stress and anxiety. There are currently no established treatment guidelines for these conditions, but emerging research indicates that internet-based cognitive behavior therapy (ICBT) is effective in the short-term for reducing these symptoms. However, the long-term stability of effects is unknown, as well as the importance of timing of the intervention. Objective: This paper aims to investigate the stability of effects achieved by ICBT after MINOCA or TS for common mental health symptoms, as well as to compare the difference in effect between early and late initiation of the intervention. Methods: This was a parallel, multicenter, randomized controlled trial of ICBT on symptoms of stress and anxiety. The ICBT intervention was a guided, 7–9-week written program, where completion of at least 5/9 steps was considered a minimum dose. Participants were included if they reported elevated stress (Perceived stress scale, 14 items [PSS-14] > 24) or anxiety (Hospital anxiety and depression scale, anxiety subscale [HADS-A] > 7) and were randomized (1:1) to either early intervention group (EIG; immediately after randomization) or late intervention group (LIG; 10-12 weeks after randomization). Mental health outcomes were assessed using PSS-14, HADS, cardiac anxiety questionnaire [CAQ], and impact of event scale [IES]. Outcome data was collected at three follow-up time points: 10-12 (1), 20-22 (2), and 50-52 (3) weeks after randomization. This paper reports results from follow-up 2 and 3. Data was analyzed by the intention-to-treat principle, using linear mixed models, assuming data missing at random. Results: Eighty-eight participants were randomized (EIG: 45; LIG: 43). The response rate at follow-up 2 was 81% and at follow-up 3 was 73%. Eighty-seven percent completed the minimum dose in the EIG, and 51% in the LIG. Intervention effects achieved in the EIG at follow-up 1 remained stable over time. The average trend was lower levels of stress and anxiety in the EIG compared with the LIG, but no differences were statistically significant. Regarding symptoms of post-traumatic stress, the EIG reported significantly lower levels than the LIG, at follow-up 2 (-2.32, 95% CI: -4.42, -0.21, p=0.03). Conclusions: Short-term effects achieved by ICBT after MINOCA or TS for symptoms of stress and anxiety were stable over time and there was little evidence suggesting a difference in effect between early or late initiation of treatment. However, patients in the EIG completed the treatment to a higher degree and reported superior treatment effects on symptoms of post-traumatic stress. These novel findings are important for informing care after a MINOCA or TS, but more research is needed before firm conclusions can be drawn. Clinical Trial: clinicaltrials.gov NCT04178434

  • Adolescents’ Use of AI Chatbots for Social, Emotional and Mental Health Support: A Cross-Sectional, School-Based Survey

    Background: An ongoing youth mental health crisis has been linked to social technologies, with gendered and developmental differences in youth mental health outcomes related to social technology exposure. Understanding their uses, attitudes about, and experiences of AI can guide policies, clinical practice, and parental and school-based approaches. Objective: Our objective was to characterize adolescents’ uses, attitudes, and experiences regarding GenAI chatbots and examine developmental stage and gender differences. Methods: A STROBE-compliant, cross-sectional evaluation using a youth-informed survey in late 2025 was conducted with school-based adolescents (ages 11-18) across eight sociodemographically diverse campuses. Exploratory survey measures assessed adolescents’ uses, attitudes, and experiences regarding GenAI chatbots. Results: 2,458 adolescents completed the survey (53% of eligible students). Participants included 38% students in grades 6 through 8 and 62% grades 9 through 12; 49% male, 49% female, 0.6% nonbinary/gender fluid; 30% Hispanic/Latine, 21% Black/African American, 18% Asian, 16% White, 9.9% multiracial, and .8% other race. Approximately three-fourths of adolescent participants reported any use of GenAI chatbots, most commonly for socioemotional support (49.5%) and help with school difficulties (62%). The mean age of first use was 12.75 years. Usage was slightly higher among females. Around one-fifth (21%) used AI to share mental health concerns, and around 13% used it to improve their mental health, with both behaviors more commonly reported by females. Early adolescents in grades six to eight (typically ages 11-13) used AI for socioemotional support more often than middle adolescents (typically ages 14-18). Regarding general attitudes, early adolescents held more positive views of AI as a wellness tool than middle adolescents. Additionally, females reported lower general AI optimism and greater AI pessimism than males. Most adolescents (65%) believed AI helped them or a friend, though 9% reported harmful experiences related to AI chatbot use. Conclusions: More adolescents may be using AI chatbots for social, emotional, and mental health support than previously estimated, with first use occurring around age 13. Early adolescence may be an optimal window for AI literacy and mental health literacy interventions, parental conversations about AI use (including risks of sycophancy and dependency), and age-appropriate access restrictions at home and school. Clinical Trial: Not applicable

  • Accessibility and Self-Management Outcomes of Patient-Facing Digital Tools for People With Eye Disease and Vision Impairment: Scoping Review

    • This manuscript needs more reviewers

    Background: Digital self-management tools for chronic eye disease are evaluated on device or algorithm performance rather than on whether people already living with functional vision loss can reach and keep using them. Whether accessibility translates into sustained self-management, and whether clinical benefit follows, is largely untested. The role of nursing mediation in that translation is unmapped. Objective: We aimed to map where evidence sits along the pathway from disease severity to clinical outcome for patient-facing digital tools in eye care, and to test whether the apparent absence of nursing-mediation evidence is real or an artefact of database coverage. Methods: Following PRISMA-ScR, we searched PubMed, Web of Science, Scopus and CINAHL Ultimate for English-language records published 2015–2026. Eligibility required eye disease or functional vision impairment, a patient-facing digital tool, and a user-side outcome. Studies were coded on causal-chain position, tool class, vision status, outcome type and evidence grade. Results: Of 7,536 records, 97 studies were included. Evidence concentrated in one link: 85 studies (87.6%) addressed interface accessibility → self-management behaviour, against 9 and 3 for the adjacent links; 76 stopped at an accessibility or usability endpoint and only 6 reached a clinical outcome. Only 21 (22%) measured any causal edge. The nursing-mediation layer is a genuine absence within our criteria — no study treated nursing attribution as an examined variable. A severity gradient runs in the same direction but is not significant (Fisher P=.12). Conclusions: Research has documented whether users can operate these tools, leaving untested whether operation changes outcomes or whether nursing mediation makes it possible. We term this the accessibility paradox and specify four measurable breakpoints in the assess–adapt–teach–follow-up sequence, so that the zero finding is refutable by design.

  • Generative AI and Medical Manuscript Integrity

    Abstract Generative artificial intelligence is increasingly used in medical manuscript preparation. Large language models may improve readability, organization, and access to scientific communication, particularly for authors who are not native English speakers. The same technologies may also create new vulnerabilities for research integrity by generating fabricated or misused references, persuasive but unsupported interpretation, and polished manuscripts that conceal weak evidence or inadequate human intellectual contribution. This position paper analyzes generative AI in medical publishing as a challenge for artificial intelligence in medicine, because unreliable manuscripts may enter evidence synthesis, clinical guidelines, educational materials, and ultimately clinical decision making. We propose a three-group framework: honest researchers who use AI to improve communication; dishonest researchers who use AI to accelerate deception; and an intermediate group of fundamentally honest authors who may adopt shortcuts under academic or financial pressure. The central risk is not visibly artificial text but credible-looking manuscripts whose form imitates scholarship while their evidentiary substance remains weak or unverified. We recommend transparent AI disclosure, manual reference verification, data availability, reviewer education, editorial screening for integrity risks, and reform of promotion systems that reward publication volume over scientific quality.

  • Applications, Predictive Models, and Clinical Integration of Artificial Intelligence in Autologous Breast Reconstruction: Scoping Review

    • This manuscript needs more reviewers

    Background: Artificial intelligence has undergone rapid development in recent years and is increasingly integrated into various medical specialties, including plastic and reconstructive surgery. However, a comprehensive mapping of AI applications specific to autologous breast reconstruction is required. Objective: This scoping review aims to systematically map the current literature regarding artificial intelligence applications in autologous breast reconstruction. Methods: A scoping review was conducted using the MEDLINE, Scopus and Embase databases. Full-text original articles investigating the use of AI in autologous breast reconstruction published between January 1, 2020, to June 30, 2026, were included. Abstracts, case reports, animal studies and studies evaluating non-autologous breast reconstruction exclusively were excluded. Thirty studies met the final eligibility criteria. Results: Among the 30 included studies, AI was predominantly applied in the preoperative phase, with emerging applications in intraoperative and postoperative care. Preoperatively, AI was utilized for risk and complication prediction, patient-reported outcomes, patient satisfaction modeling, automated imaging analysis, and interactive patient counseling. AI was applied intraoperatively to answer clinical queries and quantify perforators in DIEP flap surgery. Postoperatively, deep learning models were applied to postoperative assessment and patient-centered outcome evaluations. Conclusions: AI is currently applied extensively in the preoperative phase of autologous breast reconstruction for predictive modelling, imaging automation and patient counseling. Intraoperative and postoperative applications are equally emerging. Further studies are required to validate these models on larger, diverse datasets to ensure safety and accuracy before routine clinical adoption.

  • Associations Among Social Media Use Motivations, Body Appreciation, and Selfie Engagement in Chinese Women Aged 18 to 35 Years Using Rednote (Xiaohongshu): Cross-Sectional Questionnaire Study

    • This manuscript needs more reviewers

    Background: Selfies combine several appearance-related activities into a single routine: taking photographs, choosing among them, editing selected images, and deciding whether to post. Research has linked some of these activities, particularly comparison and editing, to poorer body image, while posting and supportive feedback can have different associations. Less is known about how the reasons people use a platform relate to this routine, or where body appreciation fits within those relationships on Chinese social media. Objective: This study examined the direct and indirect associations between four UGT-derived motivations for social media use (information seeking, social interaction, entertainment, and status seeking), body appreciation, and selfie engagement among young Chinese women who use Rednote (Xiaohongshu). Methods: A cross-sectional, self-administered web-based survey was distributed via Rednote and WeChat and hosted on Wenjuanxing in June 2026, following a May 2026 pilot (n=72). Eligible respondents were Chinese women aged 18 to 35 years who were active Rednote users. Of 567 submissions, 21 failed eligibility screening, and 94 were classified as straight-lined in the study's cleaning record, leaving 452 analyzable responses. Motivations were measured with a 12-item, 4-dimension UGT scale; body appreciation with the Body Appreciation Scale-2; and selfie engagement with the 15-item Selfie Engagement Scale, modelled as a reflective-formative higher-order construct comprising preoccupation, selection, editing, and posting. Partial least squares structural equation modelling used 5000 bootstrap resamples. Results: The measurement model met the prespecified criteria (Cronbach αlpha, 0.707-0.917; average variance extracted 0.574-0.697; all heterotrait-monotrait ratios <0.90). All four motivations were positively associated with body appreciation, jointly explaining 65.4% of its variance: social interaction (β=.310, P<.001), information seeking (β=.300, P<.001), entertainment (β=.233, P<.001), and status seeking (β=.124, P=.01). Three were directly associated with selfie engagement: information seeking (β=.222, P<.001), status seeking (β=.221, P<.001), and social interaction (β=.139, P=.02). Entertainment was not (β=.068, P=.07). Body appreciation had the largest standardized coefficient in the selfie engagement equation (β=.291, P<.001; R²=0.623). All four statistical indirect effects through body appreciation were significant (β=.036-.090; all bias-corrected 95% CIs excluded zero). Entertainment showed an indirect-only association; the other three were complementary Conclusions: In the fitted model, body appreciation accounted for a modest part of the association between each motivation and selfie engagement. Entertainment showed an indirect association only. The positive association between status seeking and body appreciation resembles an earlier finding from Korea, but culture was not measured and should not be treated as the explanation. The results support closer study of positive and negative body image processes together, using longitudinal data and measures that distinguish selfie selection, editing, and posting Clinical Trial: Not Applicable

  • Evaluating AI-Generated Physical Activity Recommendations in Primary Care: Qualitative Framework Development Study

    Background: In primary care, artificial intelligence (AI) is increasingly being developed to support personalised physical activity recommendations for people living with long-term conditions by integrating clinical, behavioural and contextual information. As integration occurs, patients, carers and healthcare professionals must decide whether AI-generated recommendations should influence clinical care. Existing qualitative studies describe stakeholder perspectives but provide limited explanation of the reasoning processes underpinning these decisions. Objective: This study aimed to develop a stakeholder-informed conceptual framework describing how people living with long-term conditions, carers and healthcare professionals evaluate AI-generated physical activity recommendations in primary care. Methods: Framework analysis was used to explore how stakeholders evaluated AI-generated physical activity recommendations. We conducted a secondary analysis of transcripts from four focus groups, which were inductively coded to identify participants’ perceptions of AI-generated recommendations. Through iterative researcher–Large Language Model dialogue, codes were synthesised into higher-order concepts across stakeholder groups. These concepts were subsequently developed into an explanatory framework describing the reasoning processes underpinning stakeholders’ evaluation of AI-generated physical activity recommendations in primary care. Finally, concepts from the COM-B model were mapped onto the framework to examine how capability, opportunity and motivation may influence responses to, and engagement with, AI-generated physical activity recommendations. Results: Stakeholders evaluated AI-generated physical activity recommendations through four interacting constructs. Decisions depended on confidence in the recommendation (construct one = credibility), its alignment with individual health and circumstances (construct two = personal fit), appropriate future clinician involvement in sustaining recommended physical activity (construct three = supported use), with any behaviour change underpinned by the individual’s capacity, motivation and opportunity to act on recommendations (construct four = behavioural appraisal). Rather than relying on a single consideration, participants integrated these constructs before deciding whether AI-generated recommendations should influence care. Conclusions: The stakeholder-informed conceptual framework developed as part of this study demonstrates that successful implementation of AI-enabled physical activity recommendations depends on how recommendations are evaluated within clinical decision-making, rather than on algorithmic performance alone. By identifying the reasoning processes used by patients, carers and healthcare professionals, and considering personal behaviour change barriers and facilitators, the framework provides practical guidance for designing AI systems that support shared decision-making, facilitate clinical implementation and improve the acceptability of AI-generated recommendations in primary care. The study also demonstrates the potential of researcher-directed LLM-assisted secondary qualitative analysis to support transparent conceptual framework development while maintaining researcher oversight.

  • Digital Engagement among Culturally and Linguistically Diverse Population: A qualitative study

    Background: Digital engagement among adults can impact healthcare outcomes and patient engagement; however there are barriers to utilizing digital tools. Objective: This study explored those barriers and facilitators to digital engagement among English and Spanish speaking adults in Houston, Texas. Methods: We conducted 4 focus groups, 2 in English and 2 in Spanish. Results: In this paper authors present the themes that emerged among them including, (1) Digital Literacy and Technology, (2) Trust, Privacy, and Security Concerns, (3) Support System and Workarounds, (4) Usability and Accessibility, (5) Communication Preferences, and (6) Digital Health Engagement. Results suggest participants desire simplicity, quick access, and easy to understand digital tools. Participants expressed familiarity with utilizing digital tools and accepted how advanced technology was a part of the world today. Participants shared frustrations and fears when providing health data to their providers. They also provided insights into educational information that might help facilitate utilization of digital tools. Conclusions: Findings from our study could inform the importance of customizing and personalizing digital tools for culturally and linguistically diverse populations served within hospital systems to optimize their adaptation and effectiveness.

  • A Rapid, Reproducible AI-Assisted Workflow for Generating Course-Aligned Anki Decks in Biomedical Sciences Education: A Tutorial

    • This manuscript needs more reviewers

    Background: The Marian University Biomedical Science (BMS) Master’s program is an intensive graduate-level program designed to mimic the rigor of the first year of medical school. This requires students to implement self-directed strategies to manage cognitive load and support retention. Anki is a digital flashcard program that enables spaced repetition and active recall, but manual flashcard generation is time consuming. Here we describe a reproducible workflow using ChatGPT 5.2 to generate Anki decks aligned to course objectives and lecture content. Objective: To describe and demonstrate a reproducible AI-assisted workflow for rapidly generating course-aligned Anki decks from instructor-provided lecture materials in medical sciences education curriculum. Methods: Lecture slides were downloaded as editable PowerPoint files and modified during class to remove identifying information, delete non-testable information, clarify key points, and condense extra content. After lecture, edited slides were exported as a PDF and provided to a ChatGPT 5.2 Thinking session with a lecture-specific prompt. ChatGPT generated two import-ready CSV files per lecture, which were imported into Anki. Results: Across six completed examination blocks, 17,222 total Anki cards were generated from BMS coursework. Mean deck size was 106 cards per lecture. Cards comprised 56% Basic (9,644) and 44% Cloze (7,578). Deck generation required ~3-5 minutes per lecture, and >99% of cards were retained in the decks with only minor edits for clarity or focus Conclusions: This AI-assisted workflow functions as a scalable learner-support tool to rapidly produce high-yield, objective-aligned Anki decks in an accelerated graduate curriculum, enabling immediate post-lecture studying and recall practice customizable to course content and individual student needs. To our knowledge, this represents one of the largest implementations of AI-assisted flashcard generation within medical curriculum.

  • The Effect of Information Summarisation and Visualisation Versus Access to More Notes on Physicians’ Case Summaries and Care Plans: A Three-Arm Randomised Controlled Trial

    • This manuscript needs more reviewers

    Background: Physicians caring for patients with complex care needs often need to review large volumes of fragmented clinical information from several sources. Whether better access to more clinical notes is sufficient, or whether information must also be summarised and visualised to support clinical work, remains uncertain. Objective: The aim was to determine whether a digital tool integrating clinical notes from several sources with elements that summarise and visualise information (“DigiTeam”) improves the quality of physicians’ written case summaries and care plans for complex patient cases, and whether any improvement reflects summarising and visualising information, access to several sources, or both. Methods: A three-arm, parallel-group, superiority randomised controlled trial among physicians in Norway. Participants were randomised to one of three digital versions: access to clinical notes from a single-source, access to clinical notes from several sources, or DigiTeam, which combined notes from several sources with summaries and visualisations. Participants reviewed three real-world pseudonymised patient cases with complex care needs and wrote a patient summary and care plan for each case. The primary outcome was the quality of the written summaries and care plans, assessed against predefined expert-derived criteria. Results: Of 93 randomised physicians, 90 contributed data. Access to several sources did not improve overall quality compared with single-source access (mean difference 0.5, 95% CI −2.3 to 3.3; p=0.970). DigiTeam produced higher overall quality scores than multi-sources access without summarising and visualising elements (mean difference 6.49, 95% CI 3.2 to 9.8; p<0.001) and than single-source access (mean difference 6.95, 95% CI 3.7 to 10.2; p<0.001). DigiTeam also improved perceived information support, while time spent and usability did not differ between groups. Conclusions: A digital tool that summarised and visualised clinical information improved the quality of physicians’ summaries and care plans without increasing time use or reducing usability. In contrast, access to more clinical notes alone did not improve quality. These findings suggest that, for complex patient cases, how information is presented matters more than how much information is available. In short, summarised and visualised information improved quality; more information alone did not. Clinical Trial: ClinicalTrials.gov, NCT07142616, registered on 21 August 2025

  • Background: Perioperative health education is important for postoperative recovery among patients with gallstones. However, traditional health education faces challenges related to insufficient continuity and personalization during perioperative information delivery. Objective: This study aimed to apply an AI agent–based health education system in perioperative health education for patients with gallstones and evaluate its effects on improving postoperative recovery outcomes. Methods: This study conducted a quasi-experimental research on patients undergoing laparoscopic cholecystectomy in the hepatobiliary surgery department of a tertiary grade A hospital in China. Patients were grouped according to their admission sequence. Patients in the control group received conventional perioperative health education, while those in the intervention group received health education through an AI-based intelligent system. The primary outcome measure was the quality of recovery 24 hours after surgery, which was evaluated using the QoR-15 scale 24 hours after surgery; the secondary outcome measures included the time of first getting out of bed, the time of first passing gas, the incidence of postoperative diarrhea, the length of hospital stay, and the evaluation of system usability. Results: The health education system based on AI agents for patients with gallbladder stones during the perioperative period can improve the postoperative recovery quality of patients within 24 hours (P < 0.05), and the time for the first standing up and the first defecation of patients are both shortened (P < 0.05); however, there were no differences between the groups in terms of the incidence of postoperative diarrhea and the length of hospital stay (P > 0.05). Conclusions: The AI agent health education system provides patients with a health education tool that is readily accessible, provides continuous responses, and offers detailed guidance. It can serve as an auxiliary tool for nursing teams to provide health education and may facilitate patients’ postoperative recovery, providing a direction for the application of artificial intelligence technology in perioperative health education.

  • Background: Prospective payment systems (PPS) impose profound financial pressures on Intensive Care Units (ICUs), often fueling administrative anxiety that treating unavoidable high-acuity cases inevitably triggers catastrophic deficits. However, empirical evidence disentangling non-linear policy safety nets from actual structural inefficiencies remains scarce. Objective: To bridge this gap, this study leverages Explainable Artificial Intelligence (XAI) to decode the complex financial risk matrix under China's Diagnosis-Intervention Packet (DIP) system. Methods: We conducted a retrospective analysis of 5,258 ICU episodes at a major tertiary teaching hospital in China, utilizing deterministically linked clinical and financial datasets. Four algorithms were benchmarked to predict continuous financial deficits. The optimal Gradient Boosting Decision Tree (GBDT) model was integrated with SHapley Additive exPlanations (SHAP) to map localized marginal impacts and extract precise empirical zero-crossing thresholds. A prespecified 'pure ICU' sub-cohort sensitivity analysis was conducted to validate structural robustness. Results: GBDT significantly outperformed traditional linear models (R2 = 0.7674). SHAP dependence analysis revealed a sophisticated policy dichotomy. A profound "double zero-crossing" in the Cost Multiplier quantified a built-in safety net, providing a maximal protective trough at 2.75x. High-acuity metrics (RW > 4.93, LOS > 20.8 days) triggered similar financial buffering. Conversely, surpassing a rigid zero-crossing threshold of 23.51% in the Diagnostic & Lab Test Ratio induced a monotonic deficit surge (the "monitoring penalty"). Crucially, the model empirically debunked traditional cost-containment dogma: medical consumables played a surprisingly marginal role, whereas strategic interventions to optimize Length of Stay (LOS) averted systemic losses. These extracted thresholds demonstrated striking stability in the sensitivity overlay. Conclusions: The DIP framework operates not as a simplistic linear penalty for high-acuity care, but as a highly engineered matrix that buffers extreme physiological crises while strictly penalizing structural inefficiencies. By debunking the traditional consumables dogma, our XAI-derived blueprint empowers hospital administrators to transition from blind cost-cutting toward strategic clinical pathway optimization. Policymakers must introduce ICU-specific risk-adjustment coefficients to mitigate unintended monitoring penalties and safeguard value-based critical care.

  • How rural Gypsy/Traveller women approach digital healthcare services as caregivers and care receivers: an interview study

    Background: Gypsy/Travellers are an ethnic group in the UK with poor health outcomes and multi-level challenges in services access, including potential digital barriers. In addition to these access and technical factors, gender also plays an important role in community healthcare in Gypsy/Traveller culture. This study explores the unique experiences of Gypsy/Traveller women in rural Scotland regarding how they engage with digital health services. Objective: To understand how Gypsy/Traveller women’s gendered roles and the use of technology shape each other in healthcare, and the implications for future services to be inclusive and supportive for this population. Methods: We recruited Gypsy/Travellers from a community organisation; participants were interviewed in small groups during community events. We drew on the concepts of intersectionality and ‘doing gender’ to explore Gypsy/Traveller women’s roles in healthcare and how technology impacted these roles. Results: Twenty-one participants (aged 16-71 years) were interviewed. Gypsy/Traveller women often manage healthcare for their family members while being healthcare recipients themselves; they have to navigate a hostile external environment due to systematic exclusion. Women are the main users of digital healthcare services, reflecting both their gendered roles and personal needs. Technology can help women juggle domestic responsibilities and provide space for them to discuss personal health, while at the same time, it may hinder relationship building with healthcare professionals and replace therapeutic in-person interactions. Conclusions: Out findings highlight how technology use in health is gendered when women are both caregivers and healthcare recipients. Technology holds the potential to support women but may also reinforce gender dynamics. To move towards a more equitable system, ethnic minority women should be included in service design; digital healthcare services meet the needs of women. The healthcare system needs to build up capacity to tackle discrimination and systematic exclusion of marginalised populations.

  • Mapping the Cancer Support Ecosystem on Reddit: A Computational Analysis of Community Connections, Supportive Care Needs, and Emotional Expression

    • This manuscript needs more reviewers

    Background: Online communities provide large-scale, real-time insights into essential concerns and needs for people affected by cancer and have become an informal extension of the broader supportive care environment. Objective: This study aimed to identify the structure, content, and emotions in cancer-related Reddit discussions using large-scale data and computational methods. Methods: We analyzed 3,206,108 cancer-related posts published across 40 subreddits from 2010 to 2025. Shared authorship and cross subreddits movement networks were constructed to examine community connections. Guided by supportive care needs frameworks and validated clinical surveys, a Large Language Model classifies posts into five key needs: informational, physical, psychological, and practical. BERTopic was used to identify topics, and a RoBERTa-based model identified seven emotions: anger, disgust, fear, joy, neutral, sadness, and surprise. Temporal trends in emotion were examined, and differences across supportive care needs were assessed using chi-square tests and Cramér’s V. Results: Discussion volume steadily increased during the study period (2010-2025) and author activity distribution was highly skewed. The structure of the network characterized that r/cancer served as a central hub linking general, disease-specific, and support-oriented communities. The most prominent needs were informational (n=1,467,967, 45.79%), followed by psychological (n=584,306, 18.22%), physical (n= 508,149, 15.85%), social (n=331,953, 10.35%), and practical needs (n=104,722, 3.27%). Emotional expression changed over time, with sadness, anger, disgust, and joy decreasing, while fear and surprise increased. Emotion patterns varied across need discussions. Although neutral emotion was common overall, informational needs were characterized by fear and surprise; physical needs by disgust and sadness; psychological and social needs by joy and anger; and practical needs by sadness and anger. Conclusions: This study suggests that online cancer discussions capture both the content of unmet needs and the emotional context in which they are experienced. These results may inform patient-centered health communication, digital supportive care interventions, and future research using social media data to monitor cancer-related needs across the cancer continuum.

  • Evaluating Human–Large Language Model Interaction in Clinical Prediction Rule Calculation: Development and Validation Study

    • This manuscript needs more reviewers

    Background: Clinical prediction rules (CPRs) are widely used in medicine, including for pediatric injury care. For retrospective validation, predictor variable extraction must be accurate and reproducible to avoid bias. However, extracting predictor variables from electronic health record (EHR) clinical notes is labor–intensive, difficult to scale, and prone to reviewer error. Objective: We aimed to evaluate a large language model (LLM) for its ability to calculate risk categories and extract predictor variables from emergency department (ED) physician notes in the EHR for four pediatric trauma CPRs. We also sought to characterize human–LLM interaction in clinical data extraction. Methods: We conducted a cross–sectional study of 750 de–identified pediatric emergency department visits in order to assess four published trauma CPRs in children: traumatic brain injury (TBI) in children younger than 2 years (TBI in younger), TBI in children older than 2 years (TBI in older), cervical spine injury (CSI), and intra–abdominal injury (IAI). The LLM performed zero–shot extraction using prespecified prompts (development cohort; n=300 visits), and performance was compared with expert reviewer extraction methods (validation cohort; n=450 visits). The primary outcome was LLM performance for CPR calculation and predictor variable extraction, including non–inferiority to expert reviewers using a prespecified 7.5% margin. Secondary outcomes examined the role of human–LLM interaction in reference standard creation. These included the number of expert reviewer extraction errors identified by the LLM, the resulting changes in children's risk categories, and the percentage decrease in extractions requiring full physician review. Results: The LLM’s risk category calculation accuracy was 91% (95% CI 85-96) for TBI in younger children, 97% (95% CI 94-100) for TBI in older, 95% (95% CI 91-98) for CSI, and 98% (95% CI 97-100) for IAI. The LLM’s risk category calculation accuracy and sensitivity were non–inferior to expert reviewers for TBI in older children, CSI, and IAI. LLM data extraction accuracy and sensitivity were also non–inferior to expert reviewers for 25 of 27 (93%) predictor variables. During reference standard creation, the LLM identified 34 human extraction errors (20% of corrections made), updated risk categories for 11 children (2.4%), and reduced physician review burden by 72%. Conclusions: An interactive human–LLM pipeline achieved expert–level performance for pediatric trauma CPR calculation and predictor variable extraction. Iterative prompt development improved reference standard quality and reduced physician review burden. These results also support a framework for human–LLM interaction spanning autonomous extraction, supervised screening, error identification, and consensus building. Our findings suggest that clinical data extraction is best viewed as an interactive process in which experts and LLMs contribute complementary strengths across multiple stages of extraction.

  • Background: Delayed Code Stroke activation contributes to worse outcomes in acute stroke. Emergency Department (ED) triage notes contain free-text clinical information that could enable automated, real-time pathway activation. We evaluated the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting Code Stroke criteria from ED triage notes. Objective: The aim of this study was to evaluate the diagnostic accuracy of a multi-pass large language model (LLM) pipeline for identifying patients meeting institutional Code Stroke criteria from ED triage notes and compare a single prompt to chain prompt performance. Methods: A retrospective cross-sectional study was conducted at Monash Medical Centre, Melbourne, Australia. De-identified triage notes from 3,023 ED presentations over a one-month period (September–October 2023) were analyzed. The pipeline applied sequential passes for translation, stroke symptom identification, mimic exclusion, baseline functional status, temporal window classification, and symptom resolution. Six locally deployed language models were evaluated. Performance was assessed against two reference standards: neurologist-labeled diagnosis and documented ED Code Stroke activation. Primary outcomes were sensitivity and specificity; secondary outcomes included PPV, NPV, and Gwet's AC1. Reliability of the neurologist reference standard was assessed by blinded independent re-review of a stratified random sample of 200 presentations by a second neurologist. Results: Of 3,023 presentations, 136 were neurologist-labeled positive. Agreement between the primary and a blinded second neurologist on a 200-note reliability sub-sample was almost perfect (raw agreement 95.0%, Cohen's κ 0.900, 95% CI 0.838–0.959). The cohort included 140 ED Code Stroke activations (median age 69, IQR 56–81 years), of whom 83 (59.2%) had confirmed stroke diagnosis. Sixteen patients (11.4%) underwent endovascular clot retrieval and 4 (2.9%) received thrombolysis. The best-performing model (Qwen 2.5 14B) achieved sensitivity 0.890 (95% CI 0.826–0.932), specificity 0.993 (0.989–0.996), PPV 0.858 (0.791–0.906), and NPV 0.995 (0.991–0.997). Pairwise McNemar testing demonstrated statistically superior overall accuracy for Qwen 2.5 14B over Llama 3.1 8B, Phi-4 14B, and Mistral 3 14B (all p<0.001 after Holm correction), with no significant difference detected versus Nemotron-Nano-12B-v2 or Qwen 3 14B. Conclusions: A locally deployed language model demonstrates acceptable sensitivity and specificity for automated Code Stroke identification from free-text triage notes. Performance was comparable across the two best models, suggesting that capable open-weight models in this parameter range may be sufficient to proceed with ongoing internal testing and external validation. The pipeline operates without internet connectivity or model retraining on patient data, supporting feasibility for real-world ED integration. Clinical Trial: ACTRN12626000596303

  • Quality of Information About Attention-Deficit/Hyperactivity Disorder on Social Media: a Scoping Review

    • This manuscript needs more reviewers

    Background: There are 5.66 billion social media user identities worldwide, and content about ADHD is popular, with 6.1 million and 5 million videos under the hashtag “#ADHD” on Instagram and TikTok respectively. Given misinformation is common on social media, it is important to establish the quality and accuracy of information about ADHD on social media. A scoping review was chosen to explore the scope and nature of research on this topic, and to facilitate comparison across different social media sites and different research methods. Objective: This review examined the methods used by researchers to assess the quality of ADHD-related information on social media. Specifically, it aimed to identify the methodologies researchers used to assess the quality of information, synthesise the existing evidence, and identify areas and issues requiring more research. Methods: The scoping review was conducted in accordance with the PRISMA extension for scoping reviews (PRISMA-ScR). Three databases were searched: Medline, Embase, and PsycInfo. The search was limited to English language and papers published from 2009 onwards. The search strategy included the names of major social media platforms. Studies were eligible if they evaluated the quality of ADHD-related information available on social media. Ten studies met the inclusion criteria. Of these, 9 were conducted in high-income countries. Data was extracted independently by 2 reviewers who subsequently collaborated to synthesize the findings and identify the main themes and conclusions. Results: Four studies analysed TikTok, 1 study analysed both TikTok and Instagram, 4 studies analysed YouTube, and 1 study analysed 5 internet forums. Number of videos/posts analysed ranged from 45 to 159. Considerable variation was observed in the methods used to assess information quality. Four studies compared content to either the diagnostic criteria given in the DSM or a derivative of the DSM (ASRS-v1.1). Three studies used different generalised quality assessment tools designed for assessing health information. Four studies categorised content as “useful” or “misleading” using study-specific criteria. Despite methodological differences, most studies concluded that ADHD-related information on social media was of poor quality. Conclusions: The heterogeneity of methods used to assess quality limits the ability to draw broad conclusions about the quality of ADHD-related information available. Future research should adopt clearly defined quality assessment tools to improve consistency and comparability. Specifically, all reviewed studies had a limited sample size and 90% were conducted in high income countries. Moreover, the role of social media algorithms in presenting different content to different groups of people was not explored. Future research should include larger, more diverse samples of social media content and investigate how algorithm-driven content and geographical context affect the quality of information users encounter.

  • Promoting mental health using a mobile health application in adolescents with a chronic somatic illness: A randomised waitlist controlled trial

    • This manuscript needs more reviewers

    Background: Over 17% of children in the Netherlands live with a chronic illness. They are at risk of developing mental health problems, affecting quality of life and health outcomes. Objective: This study evaluates a preventive, transdiagnostic mobile health intervention, the Grow It! app, aiming to enhance overall well being by creating emotional insight and reflection, whilst encouraging the use of adaptive coping strategies. Methods: Using a randomised waitlist controlled design, adolescents aged 10–18 years with a chronic somatic illness were recruited through a university medical hospital and randomised to a waitlist control group or an intervention group that used the Grow It! app. Follow up assessments were conducted immediately after the four week intervention period and three months later. After the final follow-up, the waitlist control group also used the Grow It! app and completed questionnaires immediately after and three months later. Primary outcome measures were symptoms of anxiety and depression. Data were analysed with linear mixed models on an intention-to-treat basis. Results: Between August 2021 and September 2023, 208 participants were enrolled (Mage=13.96, SD=2.20, 57.69% female), with 105 randomised to the Grow It! intervention group and 103 to the waitlist control group. At the overall group level, symptoms of anxiety (χ2(3, 497.01)=2.08, p=.1021) and depression (χ²(3, 497.01)=0.71, p=.5459) remained stable over time, indicating no significant change across the full sample. However, among adolescents who used the Grow It! app, adolescents with elevated anxiety/depressive at baseline demonstrated the greatest benefit, showing larger symptom reductions than those who initially reported low symptom levels. Conclusions: These findings suggest that the Grow It! app may offer meaningful benefits for adolescents with a chronic somatic illness experiencing elevated mood and anxiety symptoms, highlighting its potential value as a preventive and supportive tool in paediatric healthcare. Clinical Trial: International Standard Randomized Controlled Trial Number (ISRCTN) 17883961; https://doi.org/10.1186/ISRCTN17883961

  • Understanding lichen sclerosus through online narratives: A qualitative analysis

    • This manuscript needs more reviewers

    Background: Lichen sclerosus (LS) is a chronic inflammatory dermatosis associated with significant physical and psychosocial morbidity. Despite established management guidelines, patients frequently experience diagnostic delays, inconsistent counseling, and uncertainty regarding disease progression. Online patient communities provide a unique opportunity to examine lived experiences that may not be fully captured in clinical settings. Objective: To characterize patient experiences, psychosocial concerns, and emotional burden related to LS through analysis of discussions within an online patient support community. Methods: We conducted a cross-sectional computational analysis of publicly available discussions from Reddit’s r/lichensclerosus community between 2017 and 2025. Posts and comments were analyzed separately using topic modeling, psychosocial multi-label classification, and emotion and affect analysis. Temporal trends in discussion topics and psychosocial themes were also evaluated. Results: A total of 1,313 posts and 659 comments were included. Discussion centered around illness uncertainty and fear of disease progression, accounting for 64.1% of posts and exhibiting the highest levels of emotional distress. Treatment-related concerns, particularly topical corticosteroid use, represented the second most common theme (21.2%). Psychosocial burden was substantial, with support-seeking behaviors identified in 86.97% of posts, alongside frequent discussion of sexual dysfunction and body image disturbance. Clinician mistrust demonstrated the strongest association with psychosocial burden across both posts and comments. Posts were characterized primarily by first-person narratives of distress and uncertainty, whereas comments focused on reassurance, shared experiences, and practical management strategies. Emotional burden was consistently higher in posts than comments, suggesting a potential buffering effect of peer support and community engagement. Temporal analyses demonstrated minimal changes in discussion patterns over the study period. Conclusions: Online communities serve as important spaces for validation, knowledge sharing, and informal care navigation among individuals with LS. These findings suggest that clinic-based perspectives may not fully capture the daily lived experiences and psychosocial challenges associated with the disease. Improved patient-centered education, clinician-patient communication, and integration of psychosocial support into LS care may better address patient needs. Leveraging real-world patient discourse offers valuable insight into patient priorities and may help inform more responsive and empathetic models of care. Clinical Trial: Not applicable.

  • Robotic versus Laparoscopic Surgery for Colonic Diseases: An Overview of Systematic Reviews and Meta-Analyses

    • This manuscript needs more reviewers

    Background: Colonic diseases frequently require minimally invasive procedures. Laparoscopy has become the standard minimally invasive approach but is limited by 2D vision, magnified hand tremors and poor ergonomics. Robotic surgery was introduced to overcome these shortcomings, yet it suffers from longer operation time and higher costs. Multiple meta-analyses comparing the two techniques have yielded inconsistent perioperative and oncological findings, with narrow disease coverage and limited generalizability. This overview synthesizes current evidence to inform clinical decision-making. Objective: The aim of this overview of systematic reviews was to systematically evaluate the advantages and perioperative outcomes of robotic surgery compared with laparoscopic surgery for colonic diseases, and to provide evidence-based guidance for clinical decision-making. Methods: PubMed, Embase, and Web of Science were searched from inception to January 1, 2026 for systematic reviews and meta-analyses comparing robotic and laparoscopic colonic surgery. After screening against predefined inclusion and exclusion criteria, the PRISMA 2020, AMSTAR 2.0, and GRADE tools were used to assess the reporting quality, methodological quality, and level of evidence of the included reviews. Outcomes including conversion to open surgery, operative time, blood loss, complications, length of hospital stay, oncological outcomes, medical costs, and mortality were analyzed.This overview was registered on PROSPERO (CRD420261415344). Results: Twenty-four English-language meta-analyses were included. Quality assessment showed that most reviews had some reporting deficiencies; none provided a list of excluded studies or funding information. The overall evidence was predominantly low to very low. Outcome analysis indicated that robotic surgery was associated with lower conversion rate to open surgery, less intraoperative blood loss, shorter postoperative hospital stay, and shorter time to first flatus. However, operative time was significantly longer for robotic surgery. No significant differences were found between the two groups for anastomotic leakage, incisional infection, postoperative ileus, number of lymph nodes harvested, or all-cause mortality in most studies. Robotic surgery incurred higher medical costs. Conclusions: Robotic colonic surgery offers superior minimally invasive benefits and better postoperative recovery, making it suitable for patients with technically demanding conditions. Laparoscopic surgery has shorter operative time and lower costs, making it more appropriate for routine colonic lesions and resource-limited settings. The overall quality of the available evidence is limited, and more high-quality, large-sample randomized controlled trials are needed.

  • Redesigning the Optical Viewing Environment for Children’s Digital Learning With a Foldable Distant-Image Display: Randomized Controlled Trial

    • This manuscript needs more reviewers

    Background: As digital and AI-enabled learning become embedded in children’s education, the relevant health question is not only how much screen time children accumulate, but whether the optical viewing environment can be redesigned. Conventional tablets are generally viewed at short distances and require sustained accommodation and convergence. Earlier distant-image systems were predominantly fixed desktop devices, limiting storage, portability, and deployment in homes and schools. A foldable display that preserves distant-image viewing could alter the optical exposure associated with digital learning without changing course content. Objective: To compare short-term ocular responses after an identical 40-minute digital learning task delivered through an S1 foldable distant-image display or a conventional tablet in children. Methods: This single-center, prospective, parallel-group randomized controlled trial enrolled 48 children aged 6 to 16 years and allocated them in a 1:2 ratio to a conventional tablet group (n=16) or an S1 group (n=32). Both groups viewed the same online course and video content for 40 minutes. The tablet was viewed at approximately 40-50 cm; the S1 presented the visual content as a virtual image at approximately 6 m. The primary outcome was the change in noncycloplegic spherical equivalent (SE). Key secondary outcomes were changes in subfoveal choroidal thickness and a 1-to-5 visual fatigue composite score adapted from 16 Computer Vision Syndrome Questionnaire symptom domains. Exploratory outcomes were logMAR visual acuity and short-term axial length. Baseline-adjusted linear models included age and used HC3 heteroskedasticity-consistent standard errors. Results: The mean ages were 11.1 (SD 2.6) years in the tablet group and 10.2 (SD 2.2) years in the S1 group. Mean SE changed by -0.27 D after tablet viewing and +0.10 D after S1 viewing; the adjusted between-group difference was +0.38 D (95% CI 0.27-0.49; P<.001). Mean choroidal thickness decreased by 11.1 μm in the tablet group and increased by 11.2 μm in the S1 group; the adjusted difference was +22.0 μm (95% CI 14.8-29.2; P<.001). Visual fatigue increased by 0.69 points in the tablet group and decreased by 0.22 points in the S1 group; the adjusted difference was -0.62 points (95% CI -0.87 to -0.38; P<.001). The adjusted differences in logMAR visual acuity and axial length were -0.049 (95% CI -0.072 to -0.026; P<.001) and -0.012 mm (95% CI -0.023 to -0.001; P=.041), respectively. No device-related adverse events occurred. Conclusions: During a single 40-minute digital learning task, the foldable distant-image display produced more favorable short-term SE, choroidal, and visual-fatigue responses than a conventional near-viewed tablet. These findings provide randomized mechanistic evidence that desktop distant-image optics can be translated into a foldable display modality for viewing-centered digital learning. They do not establish long-term myopia prevention, control of axial elongation, or full functional equivalence to a tablet computer. Clinical Trial: Chinese Clinical Trial Registry, ChiCTR2100047059; registered June 7, 2021;

  • Decoding the related comments of emerging anti obesity drugs on China social media: Assessing the association between weight‑management policy and public discourse

    • This manuscript needs more reviewers

    Background: Obesity is a critical global public health crisis with high prevalence in China, and glucagon-like peptide-1 receptor agonists (GLP-1 RAs) have emerged as novel weight-management agents. In 2024, China implemented official weight-management policies, making it essential to analyze public discourse about anti-obesity drugs on Chinese social media to support rational promotion and public health practice. Objective: This study aimed to analyze public discussions regarding emerging anti-obesity drugs on the Chinese social media platform Weibo, and to evaluate the changes in public discourse patterns following the implementation of China's 2024 weight-management policy. Methods: This cross-sectional study analyzed 25,733 valid Weibo posts about novel anti-obesity drugs from December 2017 to December 2025, using Latent Dirichlet allocation for topic modeling and interrupted time-series analysis to evaluate changes in public discussion trends before and after the policy. Results: We identified three major themes—Medication Usage Experience, Pharmaceutical Development Dynamics, and Medication Science Communication—and 11 subtopics. Most posts came from general users and focused on semaglutide. After the policy, the proportion of Medication Usage Experience discussions decreased from 11.76% to 7.10%, the proportion of Pharmaceutical Development Dynamics remained relatively stable (from 50.51% to 50.01%), while the proportion of Medication Science Communication increased from 37.73% to 42.89%. Interrupted time-series analysis further revealed that Medication Science Communication declined sharply (coefficient = −29.16; 95% CI: −41.90 to −16.42; P < .01) and then rose significantly (coefficient = 0.05; 95% CI: 0.04 to 0.07; P < .01). Conclusions: After policy intervention, discussion initially decreased but subsequently showed a relative increase in science-oriented content, suggesting a potential shift in public focus. These findings thus indicate public responses to obesity interventions and support scientific promotion of anti-obesity drugs and weight-management strategies.

  • Measuring Content Engagement and Associated Outcomes in a Digital Health Vaccine Intervention for Black Young Adults: A Paradata Analysis

    Background: Digital health interventions (DHIs) hold promise for improving health outcomes. Tough Talks COVID-19 (TT-C) was a culturally tailored, community-informed DHI designed to support Black young adults (YAs) in the US South in making informed COVID-19 vaccination decisions. Objective: To understand how user engagement, captured through paradata, was associated with improved vaccine attitudes in the TT-C DHI. Methods: We analyzed paradata from 299 Black YAs in Alabama, Georgia, and North Carolina who used the TT-C DHI in a randomized controlled trial. The DHI included 62 self-paced activities that were analyzed using rapid content analysis to identify the specific vaccine uptake-related domain they covered (hesitancy, confidence, conspiracy beliefs, and knowledge). Intervention engagement was defined as number of activities and time spent in the DHI. We used regression models to evaluate associations between intervention engagement and changes in vaccine-related measures after three months of use. Results: Overall engagement with the DHI was high.. Greater engagement with the DHI, both in total time and number of activities completed, was associated with larger improvements for all four vaccine outcomes at three months after download (P < 0.05). The effects of the TT-C DHI were amplified when participants engaged in activities that were directly related to the domain of interest, particularly with conspiracy beliefs, which had an effect size three times larger for domain-specific engagement (β: -0.067 [95% CI: -0.130, -0.005]) compared to domain-agnostic engagement (β: -0.019 [95% CI: -0.035, -0.002]) Conclusions: Our findings underscore the value of using paradata to assess DHI effectiveness and highlight that the type of content engaged with is associated with change. These insights can guide future DHI design and real-time adaptive interventions to maximize impact. Clinical Trial: ClinicalTrials.gov NCT05490329; https://clinicaltrials.gov/study/NCT05490329

  • Marked Public Interest in Male Breast Cancer Highlights a Potential Gap in Male-Specific Access to Breast Care Services: Results of a Nationwide Infodemiological Study in Germany

    • This manuscript needs more reviewers

    Background: Breast cancer is among the most common malignancies worldwide. Despite its high public awareness and advances in diagnosis and treatment, many patients continue to experience substantial unmet informational needs. Patients and their relatives frequently seek additional information online, making infodemiological analyses a valuable tool for identifying public concerns, knowledge gaps, and potentially underserved patient populations through online search behavior. Objective: This study aimed to characterize breast cancer–related online search behavior in Germany using infodemiological methods. Methods: Breast cancer-related online search behavior in Germany was analyzed using Google Trends to assess monthly search trends from January 2016 to December 2024 and Google Ads Keyword Planner to identify and systematically analyze associated search terms. A weighted scoring system based on search frequency was applied to further investigate treatment-related side effects and rare patient subgroups. Results: The lay term showed substantially higher search interest than the medical term indicating a clear preference for non-medical terminology. High search interest was observed for treatment-related side effects, particularly alopecia and pain. Among special patient groups, male breast cancer accounted for the largest proportion of related search queries despite its low incidence. Conclusions: Future patient information strategies should place greater emphasis on lay terminology to better reflect real-world search behavior. The remarkably high search volume related to male breast cancer is particularly notable given the generally lower health information–seeking behavior among men and may indicate barriers to access medical support. Infodemiological analyses may help identify underserved groups and support more patient-centered communication strategies.

  • Shame on Who? #FastTailedGirl Discourse on Twitter and Its Implications for Black Girls' Sexual and Reproductive Health: A Qualitative Analysis

    Background: In the United States, Black girls and women are disproportionately affected by sexually transmitted infections (STIs), HIV, and sexual violence [1-3]. The "fast-tailed girl" (FTG) narrative, a gender-specific pejorative suggesting a girl is intentionally demonstrating sexual behaviors reserved for adults, has persisted in Black communities for generations and intersects with broader patterns of adultification bias that deny Black girls the protections of childhood [4-6]. Social media platforms have become significant sites where this harmful narrative is both perpetuated and contested [7,8]. Objective: This study examines Twitter discourse surrounding the #FastTailedGirl hashtag to understand how this narrative manifests online, its implications for Black women's and girls' sexual development and health, and how community members resist this harmful trope. Methods: Using Brandwatch, we retrieved 34,485 English-language tweets containing #fasttailedgirl and related permutations posted between 2011 and 2021, with retweets excluded. Tweets identified as originating outside the United States were excluded. Tweets with no country code available were retained because missing geographic metadata do not indicate that a tweet originated outside the United States. This yielded an eligible dataset of 33,099 tweets, from which a 20% random sample (N=6,899) was selected for qualitative analysis. We analyzed this sample using reflexive thematic analysis informed by intersectionality and reproductive justice frameworks [9,10]. After tweets coded as irrelevant or lacking sufficient context (code 99; n=678) were removed, the final thematic analytic sample included 6,221 tweets. Results: Two overarching themes emerged: (1) a culture of acceptance, encompassing blaming and shaming of Black girls, sexualization and predatory behavior, denial of childhood, generational transmission, and long-term consequences including under-reporting of sexual assault; and (2) resistance through provocation and inquiry, including myth-busting, activism and advocacy, and reclaiming self and body autonomy. Conclusions: The FTG narrative represents a significant public health concern that may reflect and reinforce conditions shaping sexual and reproductive health inequities among Black women and girls. Findings point to an urgent need for culturally responsive digital and community-based interventions that counter stereotype messaging, create safe spaces for empowerment, and support healing from trauma. Clinical Trial: Not Applicable

  • The role of Artificial Intelligence in interventions to reduce alcohol consumption and smoking: A systematic review

    Background: Prevalence of alcohol consumption and smoking is high across high-income countries, and interventions to decrease either can include behaviour change delivered digitally. Artificial intelligence (AI) is a broad field encompassing various techniques, which include algorithms that learn from data to perform automated tasks without explicit human programming, and could potentially increase the effectiveness of digital behaviour change interventions through personalisation and reducing barriers to engagement. Objective: This systematic review aims to summarise the evidence for the effectiveness of AI-assisted interventions to reduce alcohol consumption and smoking. Methods: We conducted a systematic review to identify and summarise evidence from randomised controlled trials (RCTs) of AI-assisted interventions for reducing alcohol consumption or smoking. Eligible trials were RCTs that reported results of an AI-assisted public health intervention for reducing engagement with alcohol consumption or smoking in a high-income country. We searched Medline (Ovid), Embase (Ovid), Web of Science (Core collection), and Scopus for relevant trials published between 2010 and 14 January 2025. We also searched for reviews of public health interventions for alcohol consumption or smoking published between 2023 and 14 January 2025, and extracted all references from relevant reviews for screening. We conducted forward and backward citation searching on all included trials (dates of searches: July to September 2025). Screening for trials was conducted independently by two reviewers. Two reviewers independently assessed risk of bias using the Cochrane Risk of Bias 2 tool. As the included trials were heterogeneous in terms of interventions, outcomes, and timepoints, we synthesised the results narratively. Results: We included 4 trials for reducing alcohol consumption (comprising 16 reports and 4,718 randomised participants): 3 trials used apps with chatbots and reported mixed evidence, and one trial, which was effective, used a rules-based AI that tailored website content. We included 12 trials for stopping smoking (comprising 31 reports and 68,659 randomised participants): 8 trials of apps or messenger chatbots and 4 trials of recommender systems, all of which had mixed evidence. For many trials, the interventions had multiple components, of which AI-assistance was only one, meaning the effectiveness of AI-assistance specifically could not be determined. Additionally, high risks of bias across almost all trials reduced confidence in the results. There were no trials using large language models (LLMs). Conclusions: There is no strong evidence of a beneficial effect of AI-assistance in public health interventions for reducing alcohol consumption or smoking in high-income countries. Future research should better describe public health interventions that use AI and embed equity considerations into their design and analysis to ensure already disadvantaged groups are not harmed further by the adoption of AI-assisted interventions in public health. Clinical Trial: PROSPERO CRD42025642317 and CRD42025642334.

  • Impact of Virtual Reality on Pain and Anxiety During Intravenous Cannulation Procedures Among Thalassemia Participants in the UAE: A Crossover Clinical Trial

    • This manuscript needs more reviewers

    Background: Participants with β-thalassemia frequently experience repeated hospital visits and cannulations, leading to chronic pain and anxiety, which adversely affect their quality of life. Virtual Reality (VR) has shown efficacy in managing the pain and anxiety associated with medical procedures. The purpose of this study was to evaluate the effects of therapeutic VR on pain, anxiety, fatigue, boredom, and participant satisfaction during intravenous (IV) cannulation procedures Objective: This study evaluated the effectiveness of therapeutic VR in reducing pain and anxiety during IV cannulation among thalassemia participants, compared with SOC. Secondary objectives included assessing the impact of VR on fatigue, boredom, and participant satisfaction. Methods: A single-center, non-randomized crossover clinical trial was conducted at the Dubai Thalassemia Center, United Arab Emirates, between May and September 2024. Participants aged >7 years undergoing routine IV cannulation received SOC during their first visit, followed by VR-assisted cannulation during two subsequent visits. Outcomes were assessed using Visual Analogue Scale (VAS) scores for pain, anxiety, fatigue, and boredom. Physiological parameters, including heart rate and blood pressure, were also recorded. Comparisons between SOC and VR sessions were performed using paired statistical tests, with statistical significance set at P<.05 Results: A total of 115 participants completed at least one SOC session, and 111 participants completed at least one VR session. Overall, 82% of the participants were older than 18 years, and 51% were male. The mean anxiety score was significantly lower in the VR group (2.24 ± 2.6) than in the SOC group (2.92 ± 2.3; p=0.02). Similarly, the fatigue score was significantly lower in the VR group (1.67 ± 1.3) than in the SOC group (2.65 ± 2.2; p=0.01). Boredom scores (1.72±1.4 vs 2.67±2.2; P=.035) were also significantly reduced with VR. Pain scores were lower in the VR group but did not differ significantly from SOC (2.23±1.8 vs 2.89±2.1; P=.07). Participant satisfaction was numerically higher with VR (90% vs 85%; P=.237). No VR-related adverse events were reported, and the intervention was well tolerated. Conclusions: The usage of VR for the intervention is feasible, safe, and generally well-tolerated by thalassemia participants. VR effectively reduces anxiety and fatigue during routine intravenous cannulations. These findings are particularly relevant for participants who undergo lifelong, repeated procedures that cumulatively contribute to procedural distress. Clinical Trial: Clinical Trial Registration: https://clinicaltrials.gov/study/NCT07099196, identifier: NCT07099196, registered 2025-07-05. This trial was registered retrospectively.

  • Supporting Quality Improvement in Oral Healthcare Using Unstructured Patient Feedback: Taxonomy Development and Mixed Methods Evaluation Study

    Background: Natural language processing (NLP) is being increasing used to analyse unstructured patient feedback (UPF) for healthcare service quality improvement. Prior studies demonstrate the analytical potential of methods like topic modelling and sentiment analysis, but are largely descriptive or conceptual, with limited translation into tools for routine clinical practice. This leaves a critical knowledge gap between computational research and its realisation within (oral) healthcare service quality improvement, where large volumes of UPF are available but difficult to interpret at scale. Objective: Develop, implement, and evaluate a clinician-facing dashboard to translate unstructured patient feedback (UPF) into actionable insights for oral healthcare service quality improvement via a hybrid NLP Pipeline. Methods: We developed a four-stage hybrid NLP pipeline, combining topic modelling, sentiment analysis, and multi-label text classification. We then applied it to 57,794 reviews of NHS dental practices in England (2019–2024). Topic modelling using BERTopic identified 191 topics, with sentiment analysis via a fine-tuned DeBERTa model across four classes (positive, negative, neutral, and mixed). We iteratively consolidated topics into a ten-theme taxonomy through a hybrid approach integrating LLM-assisted classification with expert qualitative interpretation. The taxonomy informed a supervised multi-label text classifier, adapted for Google Maps Places reviews to assess portability, deploying it within a dashboard that processes real-time patient reviews, visualising thematic and sentiment insights. We evaluated the ten-theme taxonomy composition and dashboard useability through qualitative applied thematic analyses of reviews and ten semi-structured interviews with dental professionals. Results: Topic modelling generated 191 topics, consolidated into a ten-theme taxonomy. The sentiment classifier achieved F1=0.952 across four classes, while the multi-label theme classifier achieved micro-F1=0.765 and ROC-AUC of 0.935. Operationalised within a clinician-facing dashboard, these models enabled near real-time synthesis of patient feedback at practice level. Qualitative useability evaluation indicates the dashboard helps identify areas for service quality improvement that would otherwise be difficult to detect. Dual-axis representation of theme and sentiment enabled more nuanced interpretation, going beyond binary or single-label approaches. Overall, the dashboard helped summarise reviews for staff meetings and QI, but users tended to focus on negative feedback. Conclusions: Our study demonstrates a reproducible NLP pipeline to produce practical taxonomies both for oral healthcare and potentially other patient-feedback contexts. It addresses a critical knowledge gap in translating NLP research into clinician-facing tools for real-time service quality improvement. By integrating computational methods with domain-specific expert interpretation, we provide a way to bridge between data analysis and clinical application. Our findings highlight the need for hybrid approaches incorporating expert assessment to address error, bias, and useability. Meanwhile, our dashboard and mapping of its development pipeline offer a practical approach for embedding patient perspectives within routine digital healthcare to support data-driven, patient-centred service improvement in oral healthcare and beyond.