Journal of Medical Internet Research
The leading peer-reviewed journal for digital medicine and health and health care in the internet age.
Editor-in-Chief:
Gunther Eysenbach, MD, MPH, FACMI, Founding Editor and Publisher; Adjunct Professor, School of Health Information Science, University of Victoria, Canada Rachele Hendricks-Sturrup, DHSc, MSc, MA, FACTS, Lead Editor; Research Director of Real-World Evidence, Duke-Margolis Institute for Health Policy, Washington, DC
Impact Factor 8.2 More information about Impact Factor CiteScore 10.4 More information about CiteScore
Recent Articles

A substantial proportion of clinically relevant information remains locked in unstructured narrative documents, creating a bottleneck for clinical research, biobank annotation, registry development, and real-world evidence generation. While large language models (LLMs) enable advanced clinical text mining, adoption is constrained by concerns regarding data security, multilingual performance, and reproducibility. Manual data abstraction remains predominant for registry curation and retrospective research, despite being labor intensive, costly, and prone to variability.


Prevailing ethical oversight of data science health research concentrates on privacy, consent, bias, and fairness. These concerns are necessary but insufficient, because each presupposes an answer to a prior question that is seldom asked directly. That is, “do the targets, proxies, labels, classifications, ontologies, and population descriptors on which a current study rests still truthfully represent the persons, populations, and phenomena they are taken to describe, at the point of use rather than the point of collection?” In this viewpoint, we name that question representational veracity (RV) and develop it as a construct for upstream ethical review. Our aims are to define RV and derive the domains along which it can be assessed; to demonstrate that it asks something that measurement validity, critical data studies, and algorithmic fairness do not; and to translate it into instruments that review bodies can use. We derive 4 assessment domains of RV analytically, asking for each transition in the data journey what must remain stable for a stored artifact still to stand for what it originally stood for. The resulting domains are material provenance, informational descriptors, normative authorization, and relational community. These domains interact but do not substitute for one another. Intact provenance cannot repair a poorly chosen target, and a transparent labeling process cannot confer authorization it never had. Drawing on scholarship in quantification, classification, measurement, critical data studies, algorithmic fairness, and health AI governance, we show that a model may be accurate, reproducible, and formally fair while resting on a representation that is too thin, too unstable, or too normatively misdirected for the proposed use. We examine 4 recurrent failure modes, proxy substitution, category misassignment, label generation error, and descriptor sedimentation, anchoring each in a published case, and we present a counterpoint in which better representation reveals rather than conceals inequity. A polygenic risk score (PRS) case study illustrates all 4 domains and shows how a score can misclassify risk in the populations least represented in its derivation while its code, pipeline, and internal validation statistics remain intact. We then translate the framework into practice using 10 reviewer prompts that an editor can paste into a review form, a justification template and scoring rubric provided as appendices, a tiered model that triggers full review only for subgroup, equity, transportability, public health, or clinical implementation claims, and a graded account of what should follow an adverse finding. Our argument is that existing governance mechanisms require an upstream layer. Investigators should be asked to justify not only whether their models perform, but whether their representations are truthful enough for the claims at hand. The intended audience is investigators, informaticians, research ethics committees, institutional review boards, data access committees, funders, regulators, and journal editors.


Generative AI (GenAI) tools powered by large language models (LLMs) are increasingly used by the public to seek health information. Unlike traditional web search, these systems generate conversational responses that may alter how users assess credibility, manage uncertainty, verify information, and decide whether to consult clinicians. As GenAI becomes more embedded in everyday health information practices, a clearer synthesis of the emerging empirical evidence is needed.

Phthalates are environmental endocrine-disrupting chemicals widely used in plastics, cosmetics, food packaging, and personal care products. Women may experience frequent exposure through everyday consumer and household products. Improving phthalate-related health literacy may support informed exposure-reduction decisions; however, conventional health education provides limited opportunities for repeated, interactive, and individually tailored learning.


Generative AI (GAI) is rapidly transforming research practices, including qualitative methods in health research. While these tools offer efficiency in processing large volumes of textual data, concerns remain regarding their methodological rigor, interpretive capacity, equity, and ethical implications.

Online appointment platforms are often viewed as tools of convenience, but they may reproduce and potentially amplify inequities rooted in underlying insurance and reimbursement structures. In this commentary, we highlight implications from a study using simulated profiles of patients with statutory health insurance (SHI) and private health insurance (PHI) to conduct an internet-based audit of online appointments. Not only was the time to the first available appointment longer for SHI than PHI patients, but also the PHI profile had access to more listings with appointments, and listings offering earlier PHI than SHI appointments appeared higher in search results than listings offering the same appointment timing to both. Accounting for strengths, limitations, and opportunities for future research, the study suggests several policy implications. First, leaders should view online scheduling platforms as access infrastructure, not merely convenience tools, and evaluate them with respect to access and disparities, not only adoption and satisfaction. Second, future evaluation and work should treat access as multidimensional, assessing multiple measures of access and comparing online with other scheduling channels. Third, insurance and delivery reform should accompany actions to govern digital platforms, which can organize and display, but cannot alone eliminate, disparities rooted in reimbursement and care structures. Ultimately, the relevant question about online scheduling platforms is not only whether such systems make booking easier but also whether they narrow, preserve, or widen broader health system inequalities. Digital front doors should be judged both by how easily they open and whether patients have equitable opportunities to pass through them.

Systemic lupus erythematosus (SLE) is a multifactorial autoimmune disease influenced by genetic, epigenetic, ecological, and environmental factors, with a global prevalence of 7.7 to 13 per 100,000, and standardized mortality rates of 2.4% to 5.9%. Between 14% and 75% of patients experience psychiatric comorbidities such as anxiety and depression, which impair treatment adherence and health-related quality of life. Social media has become an important channel for patients to express health concerns and seek support. Reddit and Weibo, as mainstream platforms globally and in China, respectively, host large volumes of user-generated content; however, no prior study has examined SLE-related discourse across both cultural and platform contexts.


















