Journal of Medical Internet Research
The leading peer-reviewed journal for digital medicine and health and health care in the internet age.
Editor-in-Chief:
Gunther Eysenbach, MD, MPH, FACMI, Founding Editor and Publisher; Adjunct Professor, School of Health Information Science, University of Victoria, Canada Rachele Hendricks-Sturrup, DHSc, MSc, MA, FACTS, Lead Editor; Research Director of Real-World Evidence, Duke-Margolis Institute for Health Policy, Washington, DC
Impact Factor 8.2 More information about Impact Factor CiteScore 10.4 More information about CiteScore
Recent Articles

Large language models (LLMs) are increasingly demonstrating the potential to reach human-level performance in generating clinical summaries from patient-clinician conversations. LLMs are usually evaluated against clinical summaries that focus mainly on patients’ biology and not on their biography (eg, preferences, values, wishes, and concerns). To achieve patient-centered care, artificial intelligence clinical summarization must incorporate patient-centered domains, implemented through patient-centered summaries (PCSs).

Ecological momentary interventions (EMIs) offer a promising strategy for targeting putative mechanisms of mental health problems by delivering real-time, tailored intervention components that adapt to person, moment, and context based on data collected using ecological momentary assessment (EMA). However, most research to date focuses on effects on distal outcomes, at the person level, whereas exploration of processes at the microlevel, that is, proximal effects of EMI components on putative momentary mechanisms and outcomes, remains very limited.

Digital health has provided caregivers with access to supportive resources without space-time restrictions. Caregivers’ digital health engagement behaviors can help them track their own health and that of care recipients as well as communicate with others. While digital health tools have become more prevalent since the COVID-19 pandemic, the trend in caregiver engagement has been less explored.

A substantial proportion of clinically relevant information remains locked in unstructured narrative documents, creating a bottleneck for clinical research, biobank annotation, registry development, and real-world evidence generation. While large language models (LLMs) enable advanced clinical text mining, adoption is constrained by concerns regarding data security, multilingual performance, and reproducibility. Manual data abstraction remains predominant for registry curation and retrospective research, despite being labor intensive, costly, and prone to variability.


Prevailing ethical oversight of data science health research concentrates on privacy, consent, bias, and fairness. These concerns are necessary but insufficient, because each presupposes an answer to a prior question that is seldom asked directly. That is, “do the targets, proxies, labels, classifications, ontologies, and population descriptors on which a current study rests still truthfully represent the persons, populations, and phenomena they are taken to describe, at the point of use rather than the point of collection?” In this viewpoint, we name that question representational veracity (RV) and develop it as a construct for upstream ethical review. Our aims are to define RV and derive the domains along which it can be assessed; to demonstrate that it asks something that measurement validity, critical data studies, and algorithmic fairness do not; and to translate it into instruments that review bodies can use. We derive 4 assessment domains of RV analytically, asking for each transition in the data journey what must remain stable for a stored artifact still to stand for what it originally stood for. The resulting domains are material provenance, informational descriptors, normative authorization, and relational community. These domains interact but do not substitute for one another. Intact provenance cannot repair a poorly chosen target, and a transparent labeling process cannot confer authorization it never had. Drawing on scholarship in quantification, classification, measurement, critical data studies, algorithmic fairness, and health AI governance, we show that a model may be accurate, reproducible, and formally fair while resting on a representation that is too thin, too unstable, or too normatively misdirected for the proposed use. We examine 4 recurrent failure modes, proxy substitution, category misassignment, label generation error, and descriptor sedimentation, anchoring each in a published case, and we present a counterpoint in which better representation reveals rather than conceals inequity. A polygenic risk score (PRS) case study illustrates all 4 domains and shows how a score can misclassify risk in the populations least represented in its derivation while its code, pipeline, and internal validation statistics remain intact. We then translate the framework into practice using 10 reviewer prompts that an editor can paste into a review form, a justification template and scoring rubric provided as appendices, a tiered model that triggers full review only for subgroup, equity, transportability, public health, or clinical implementation claims, and a graded account of what should follow an adverse finding. Our argument is that existing governance mechanisms require an upstream layer. Investigators should be asked to justify not only whether their models perform, but whether their representations are truthful enough for the claims at hand. The intended audience is investigators, informaticians, research ethics committees, institutional review boards, data access committees, funders, regulators, and journal editors.


Generative AI (GenAI) tools powered by large language models (LLMs) are increasingly used by the public to seek health information. Unlike traditional web search, these systems generate conversational responses that may alter how users assess credibility, manage uncertainty, verify information, and decide whether to consult clinicians. As GenAI becomes more embedded in everyday health information practices, a clearer synthesis of the emerging empirical evidence is needed.

Phthalates are environmental endocrine-disrupting chemicals widely used in plastics, cosmetics, food packaging, and personal care products. Women may experience frequent exposure through everyday consumer and household products. Improving phthalate-related health literacy may support informed exposure-reduction decisions; however, conventional health education provides limited opportunities for repeated, interactive, and individually tailored learning.


Generative AI (GAI) is rapidly transforming research practices, including qualitative methods in health research. While these tools offer efficiency in processing large volumes of textual data, concerns remain regarding their methodological rigor, interpretive capacity, equity, and ethical implications.

















