Abstract
Background: Vascular compromise remains the leading cause of free-flap failure. AI-based monitoring and prediction tools have emerged as a promising adjunct for postoperative free flap monitoring and early detection of vascular compromise. Previous systematic reviews included limited evidence or broadly evaluated reconstructive outcomes.
Objective: This systematic review and meta-analysis assessed the diagnostic accuracy of AI for postoperative free flap monitoring and flap compromise detection.
Methods: Following PRISMA-DTA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies), PubMed, Embase, Cochrane, Web of Science, and Scopus were searched from inception to June 21, 2026. Studies developing or validating AI models for free flap monitoring or compromise prediction were included. Pooled sensitivity, specificity, area under the curve (AUC), and diagnostic odds ratio (DOR) were estimated using a hierarchical bivariate random-effects model. Risk-of-bias and applicability were assessed with Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2), and the certainty of evidence was evaluated using the GRADE (Grading of Recommendations Assessment, Development and Evaluation) approach. Prespecified subgroup analyses by modality, model-fit diagnostics, sensitivity analyses, and publication-bias testing with Deeks funnel plot were undertaken. The protocol was prospectively registered in PROSPERO (CRD420251175572).
Results: Of 2098 records identified, 18 studies met the inclusion criteria. Of these, 17 studies were included in the quantitative synthesis. Pooled analysis demonstrated an overall sensitivity of 0.83 (95% CI 0.70‐0.91; prediction interval [PI] 0.21‐0.99), specificity of 0.87 (95% CI 0.65‐0.96; PI 0.03‐1.00), AUC of 0.92 (95% CI 0.90‐0.94), and DOR 36.17 (95% CI 8.27‐158.26; PI 0.09‐14886.43). Image-based AI models demonstrated superior performance, with a sensitivity of 0.92 (95% CI 0.81‐0.97; PI 0.43‐0.99), specificity of 0.95 (95% CI 0.86‐0.98; PI 0.43‐1.00), and AUC of 0.98 (95% CI 0.96‐0.99). Nonimage-based models had sensitivity (0.69, 95% CI 0.45‐0.86; PI 0.11‐0.98) specificity (0.72, 95% CI 0.18‐0.97; PI 0.00‐1.00), and AUC (0.79, 95% CI 0.75‐0.82). For vascular compromise detection, pooled sensitivity was 0.92 (95% CI 0.81‐0.97; PI 0.44‐0.99) and specificity was 0.91 (95% CI 0.79‐0.97; PI 0.35‐1.00). Between-study heterogeneity was substantial, QUADAS-2 identified methodological concerns in several studies, and GRADE rated the certainty of evidence as low.
Conclusions: AI demonstrated favorable diagnostic performance for detecting free flap vascular compromise, particularly with image-based models. These findings support the use of AI as an adjunct to conventional postoperative flap monitoring. Nevertheless, the low certainty of evidence, substantial heterogeneity, and limited external validation indicate that prospective multicenter validation and standardized reporting are required before routine clinical implementation.
doi:10.2196/91174
Keywords
Introduction
Microsurgical free tissue transfer has become a cornerstone of modern reconstructive surgery and is routinely performed in breast, head and neck, extremity, and other complex reconstructions [,]. With reported success rates of 94%‐99%, it offers durable and functional restoration for even the most challenging defects [-]. Despite continued technical refinements and attentive perioperative care, complications remain a persistent concern []. Among these, flap failure is particularly devastating. It is most often caused by vascular compromise, including arterial insufficiency, venous congestion, and ischemia-reperfusion injury, with incidence rates reported in 10%‐30% of cases []. These events frequently necessitate reoperation, with 2%‐5% of flaps requiring surgical exploration, thereby increasing morbidity, prolonging hospitalization, escalating healthcare costs, and imposing substantial psychological stress on patients [-].
Because salvage success is highly dependent on the interval between vascular compromise and reintervention, accurate risk identification, early recognition of compromised flaps, and timely response are critical []. Risk stratification and vigilant postoperative monitoring remain essential components of high-quality microsurgical care []. However, current assessment relies heavily on clinical observation, typically performed by staff and interpreted by experienced microsurgeons [,]. These approaches are labor-intensive, require continuous bedside monitoring, and are inherently subjective, often ambiguous and challenging to interpret. Identifying key predictive factors and prioritizing high-risk patients for intensified monitoring may reduce the risk of flap failure and improve microsurgical outcomes [-].
In recent years, advances in computing and data science have accelerated the integration of AI into clinical workflows []. AI systems are capable of emulating aspects of human perception and reasoning while reducing fatigue-related errors and certain systematic biases []. Among these, convolutional neural networks, which process images through hierarchically arranged filters analogous to the visual cortex, have demonstrated utility in medical imaging tasks. These include anatomical landmark detection, vascular segmentation, and perforator mapping, where high accuracy is essential for surgical planning [,].
Within AI, machine learning, a data-driven approach that enables pattern recognition and predictive modeling, has shown promise in various domains of plastic surgery, including wound analysis, surgical planning, dermatologic diagnosis, and postoperative outcome prediction [,-]. A growing number of studies have also investigated the use of AI in predicting and assessing flap-related complications []. While these tools hold considerable promise in augmenting clinical judgment and reducing human error, the diagnostic accuracy and clinical reliability in free flap monitoring remain insufficiently established [,]. This uncertainty poses a significant barrier to widespread clinical implementation.
Two systematic reviews [,] have recently summarized the application of AI in free flap surgery. However, the available evidence remains incomplete in several respects. First, both reviews framed the outcome broadly as the prediction of postoperative complications rather than the diagnostic performance of AI for real-time flap monitoring and early recognition of vascular compromise, the specific scenario in which the interval to reexploration determines salvage []. Second, neither review distinguished image-based surveillance models from clinical variable-based risk-prediction models, although these approaches differ fundamentally in their inputs, intended point of care, and expected diagnostic behavior, and may therefore warrant separate synthesis []. Third, both searches predate a rapidly expanding body of image-based monitoring evidence, including camera-integrated deep-learning platforms, optical detection of venous congestion, and remote real-time surveillance systems [-]. These gaps underscore the need for an updated diagnostic test accuracy synthesis that reflects current clinical applications of AI and separately evaluates image-based and nonimage-based models according to their intended roles in postoperative free flap care.
Therefore, we conducted a systematic review and meta-analysis to evaluate the diagnostic performance of AI-based approaches in predicting and detecting free flap compromise following microsurgical reconstruction. We also compared the diagnostic performance of image-based and nonimage-based AI models to better characterize their role in postoperative free flap monitoring.
Methods
Overview
This systematic review and meta-analysis was registered with PROSPERO (International Prospective Register of Systematic Reviews; CRD420251175572) and conducted in accordance with the PRISMA-DTA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies) guidelines [] and the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 statement []. The study followed Transparency in the reporting of AI (TITAN) recommendations []. The literature search was reported in accordance with the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses literature search extension) extension to enhance the transparency and reproducibility of the search process. The complete search strategies for all databases and the completed PRISMA-S checklist are provided in and , respectively [].
Search Strategies
A comprehensive search was performed across PubMed, Embase, Cochrane, Web of Science, and Scopus databases for studies published up to June 21, 2026, without language restrictions (). Database searches were updated before manuscript revision to capture recently published studies. The search was designed to identify studies that developed and validated deep learning or machine learning algorithms for the diagnosis of free flap status following microsurgery, using either imaging or clinical variables. Keywords and MeSH terms included “Artificial Intelligence,” “Deep learning,” “Machine learning,” “Radiomic,” “Automated diagnosis,” “Free flap reconstruction,” “Flap monitoring,” “Flap viability,” “ischemic complications,” and “Flap failure” (). Search strategies were developed independently by the investigators and were not adapted from previous systematic reviews [,]. Backward citation searching was conducted by manually screening the reference lists of all included studies and relevant review papers. Forward citation searching was also performed to identify studies citing the included papers. No other supplementary search methods were used.

Selection Criteria
Eligibility assessment was conducted independently by two reviewers (KCH, CHL) who screened titles and abstracts of the search results, with discrepancies resolved through discussions. The inclusion criteria were: (1) original research paper evaluating the performance of AI in identifying or predicting flap compromise, regardless of model type; (2) cohort studies incorporating diagnostic analysis; (3) study populations consisting of patients who underwent microsurgical reconstruction; and (4) clearly defined diagnostic criteria for flap compromise as stated in the study with available and extractable data. Studies were excluded based on the following criteria: (1) nonoriginal study types, including systematic reviews, review papers, case reports, commentaries, and letters; (2) irrelevant research focus, such as basic science experiments; (3) studies lacking AI applications or relevance to flap-related outcomes; (4) incomplete diagnostic data, particularly when key metrics such as true positive (TP), true negative (TN), false positive (FP), and false negative (FN) could not be determined; (5) unavailable full-text articles; and (6) duplicate publications. There were no language restrictions applied in the search strategy.
Data Extraction and Management
The data extraction framework was developed with reference to the Cochrane Handbook guidance [,]. Data extraction was performed independently by two reviewers (KCH and YYC) using a predefined data extraction sheet. Extracted variables included study design, country, year of publication, authorship, patient population characteristics such as gender and age, and model-related characteristics including training and testing sample sizes, validation methods, reference standards, AI model type, and modality of flap assessment. Diagnostic performance metrics, including sensitivity, specificity, and the corresponding TP, FP, TN, and FN counts, were collected. When studies included both animal and human experiments, only human data were extracted for quantitative synthesis. Information regarding missing data handling and imputation methods was also extracted when reported.
Study independence was confirmed based on author group, region, study period, and institution, with no evidence of overlapping cohorts, and analyses were conducted at the patient level with image- or pixel-level data aggregated accordingly to avoid distortion of diagnostic estimates []. When multiple models were reported within a study, the best-performing model identified by the original authors was selected. If no preferred model was specified, the algorithm with the highest area under the curve (AUC) was chosen, with priority given to higher sensitivity because of the clinical importance of early flap compromise detection. To avoid overrepresentation of individual cohorts, only one model from each study was included in the pooled analysis. Contingency tables were reconstructed directly from reported results or supplementary data or back calculated from available performance metrics. When necessary, corresponding authors were contacted for clarification or provision of additional data.
Risk-of-Bias and Quality Assessment
Methodological quality was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool, which evaluates four domains: patient selection, index test, reference standard, and flow and timing. Each domain was rated independently by two reviewers (KCH, MJW) for risk of bias as low, high, or unclear, and the first 3 domains were also appraised for applicability concerns []. Discrepancies were adjudicated through discussion or consultation with a third senior reviewer (CHL). The certainty of evidence was assessed using the GRADE (Grading of Recommendations Assessment, Development and Evaluation) approach for diagnostic test accuracy studies [].
Diagnostic Performance and Statistical Analysis
Statistical analyses were performed using R (R Foundation for Statistical Computing) and Stata version 18.0 (StataCorp LLC) [-]. A bivariate random-effects model was applied to jointly synthesize sensitivity and specificity, generating pooled estimates with corresponding 95% CIs. Diagnostic odds ratios (DOR) were also computed. Overall diagnostic performance was visualized using summary receiver operating characteristic (SROC) curve, with the AUC reported as an indicator of discriminatory capacity, where values approaching 1.0 reflected superior accuracy []. For univariate random-effects analyses, the Sidik-Jonkman estimator was used to estimate between-study variance, and the Hartung-Knapp adjustment was applied to calculate the 95% CIs in R. Corresponding 95% prediction intervals (PIs) were calculated and displayed in the forest plots. Sensitivity and specificity were pooled as logit-transformed proportions, whereas DOR and likelihood ratios were pooled on the natural logarithmic scale. For studies containing at least one zero cell among the TP, FP, TN, or FN counts, a continuity correction of 0.5 was added to all four cells before calculation of the DOR, likelihood ratios, and their variances. No continuity correction was applied to studies without zero cells. These supplementary analyses did not replace the primary bivariate diagnostic meta-analysis []. Statistical significance was defined as P<.05.
Clinical Value Verification
The clinical utility of AI-based diagnostic approaches was further explored through Fagan plot analysis, which translated pooled diagnostic metrics into posttest probabilities, thereby enabling estimation of their clinical applicability in predicting flap compromise and related complications [,]. The pooled prevalence of flap compromise was estimated by dividing the total number of TP and FN cases by the overall sample size (TP+FN+FP+TN) across included studies. Publication bias was assessed using Deeks funnel plot asymmetry test, with significant asymmetry (P<.10) interpreted as suggestive of small-study effects or selective reporting.
Subgroup Analysis
To address potential sources of heterogeneity, prespecified subgroup analyses were conducted according to study methodology and diagnostic focus (, illustrating approaches from: first approach used perioperative clinical variables for risk prediction of flap-related outcomes, second approach applied postoperative image-based analysis to identify flap compromise through changes in flap color and surface characteristics). Studies were categorized as image-based models for real-time postoperative flap surveillance and early detection of flap compromise, or nonimage-based models using clinical and perioperative variables for risk stratification and prediction of flap perfusion-related complications. Additional subgroup analyses focused on models specifically designed to detect vascular compromise. Results were presented using SROC curves and forest plots to facilitate subgroup comparisons.
Heterogeneity and Sensitivity Analysis
Between-study heterogeneity was further quantified using the I² statistic, with values greater than 50% considered to represent substantial heterogeneity. PIs were additionally calculated to provide a clinically relevant interpretation of between-study variability []. A bivariate box plot was also generated to visually inspect the distribution of sensitivity and specificity and to identify potential outliers.
The robustness of the primary meta-analysis was interrogated through sensitivity analyses, which included assessments of goodness of fit via Q–Q plot of deviance residuals, evaluation of bivariate normality using a Mahalanobis distance-based Q–Q plot, identification of influential studies through Cook distance, and systematic detection of outliers. Outliers were defined as studies with standardized residuals exceeding 2.0 for sensitivity of FP rate, corresponding to a 99% CI (z score±2.0), or those demonstrating significant deviation from expected model fit. A conservative approach was employed to minimize inappropriate exclusions. Identified outliers were removed and the bivariate model was reestimated to determine their impact on pooled performance metrics and heterogeneity. Publication bias was assessed using Deeks funnel plot asymmetry test, with P<.10 considered suggestive of significant small-study effects.
Ethical Considerations
Ethical approval was not required because this study was a systematic review and meta-analysis based exclusively on published data.
Results
Literature Search
The process of study acquisition is illustrated in the PRISMA flow chart (). A total of 2098 studies were initially retrieved, of which 792 studies were excluded as duplicates. After screening the remaining 1306 papers by title and abstract, 1177 were excluded, leaving 129 full text studies for detailed assessment. Following eligibility evaluation, 111 studies were excluded, resulting in 18 studies [,-,-] ultimately included in this meta-analysis.

Seventeen studies included in the quantitative diagnostic accuracy meta-analysis are summarized in and . All evaluated AI for the prediction or detection of free flap compromise. One additional study [] met the eligibility criteria and is presented in the PRISMA flow diagram and ; however, it was excluded from the quantitative synthesis because insufficient data were available to reconstruct the 2×2 contingency table.
| Author, year | Country | Time period | Research type | Included patients (n) | Mean age, range (, years) | Female (%) | Reference standard | Targeted condition | Type of internal validation | External validation |
| Zhang et al, 2026 [] | China | NR | Retrospective | 7 | NR | NR | Clinical evaluation | Vascular complications | Hold-out validation: train–test split validation (4:1 split) | NR |
| He et al, 2026 [] | China | January 2018-December 2024 | Retrospective | 576 | 57 | 79.34 | Clinical evaluation | Free flap failure | 6:2:2 train: validation: test split | NR |
| Huang et al, 2026 [] | China | January 2019-June 2025 | Retrospective | 615 | NR | NR | Clinical evaluation | Venous congestion | Patient-level 10-fold cross-validation | Yes |
| Kim et al, 2026 [] | South Korea | June 2021-March 2024 | Retrospective | 131 | 58.1 | 38.2 | Clinical evaluation | Flap failure | 5-fold cross-validation | NR |
| Geovani et al, 2025 [] | Indonesia | NR | Preliminary study | NR | NR | NR | Image label: visual assessment | Flap compromise | Train: test split with validation set | NR |
| Fang et al, 2025 [] | China | January 2021-January 2024 | Retrospective | 341 | 59.41 | 43.07 | Clinical evaluation | Vascular complications | Internal split-sample validation: bootstrap internal validation | NR |
| Monarchi et al, 2025 [] | Italy | NR | Retrospective | 125 | 52.3 | 41.6 | Clinical evaluation | Flap takeback | K-fold cross-validation | NR |
| Maktabi et al, 2025 [] | Germany | NR | Retrospective | 59 | NR | NR | Clinical evaluation | Flap malperfusion | Leave-one-patient-out cross-validation (LOPOCV) | Yes |
| Oleru et al, 2025 [] | United States | January 2019-January 2024 | Retrospective | 458 | 50.5 | 52.90 | Clinical evaluation | Flap takeback due to vascular complications | 10-fold cross-validation | NR |
| Kim et al, 2024 [] | South Korea | March 2020-August 2023 | Prospective | 305 | 62 (8‐86) | 41.6 | Clinical evaluation | Vascular complications | 5-fold cross-validation, Hold-out (train-test split) validation | Yes |
| Yang et al, 2024 [] | China | January 2019-December 2021 | Retrospective | 570 | 49.5 | 28.4 | Surgical confirmation | Vascular complications | Split-sample (60:40), SMOTE | NR |
| Wang et al, 2024 [] | United States | 2012‐2019 | Retrospective | NR | 63 (55, 71) | 32 | Clinical evaluation | Reoperation | Stratified k-fold cross-validation | Yes |
| Huang et al, 2023 [] | Taiwan | January 2021-July 2021 | Prospective | 176 | NR | NR | Clinical evaluation | Vascular complications | 10-fold cross-validation | NR |
| Hsu et al, 2023 [] | Taiwan | April 2021-March 2022 | Retrospective | 642 | 55.7 (43.2‐68.2) | 24 | Clinical evaluation, surgical confirmation | Venous congestion | Random split of photograph | Yes |
| Asaad et al, 2023 [] | United States | January 2005-December 2018 | Retrospective | 4000 | NR | NR | Clinical evaluation | Total flap loss | 10-fold cross-validation | NR |
| Tighe et al, 2022 [] | United Kingdom | NR | Retrospective | 1593 | NR | 46.7 | Clinical evaluation | Complete flap failure | 10-fold cross-validation | NR |
| Shi et al, 2022 [] | China | January 2006-December 12, 2020 | Retrospective | 946 | 42 (range 13‐65) | 58.3 | Clinical evaluation | Flap failure | 5-fold cross-validation | NR |
| O’Neill et al, 2020 [] | Canada | 2009/01-2017/01 | Retrospective | 1012 | 50.9 | NR | Clinical evaluation | Total flap failure | Train:test split with 3 ratios (50:50, 60:40, 70:30) | NR |
aNR: not reported
bSMOTE: synthetic minority oversampling technique.
Most included studies were retrospective single-center cohort studies, although several investigations incorporated multicenter or externally validated datasets. presents the basic study characteristics, including design, patient population, and clinical context, while details model specifications, validation methods, and diagnostic performance metrics. Most studies were retrospective (n=15) in design, with only two prospective investigations conducted by Huang et al [] and Kim et al []. External validation was reported in five studies [,,,,]. Eight studies [-,-,,] applied image-based analysis, using photographic data to evaluate flap viability primarily based on colorimetric features. The remaining 9 studies [,-,,] developed predictive models using diverse clinical variables to identify patients at risk for flap compromise. These models integrated preoperative (eg, laboratory values, neoadjuvant therapy, and medication history), intraoperative details (eg, operative duration, ischemic time, flap selection, and laterality), and postoperative parameters (eg, hypotensive episodes) details. Demographic and clinical factors, including age, BMI, comorbidities, and primary tumor characteristics such as metastatic status, were also incorporated. Although the included studies encompassed a variety of free flap types, the anterolateral thigh (ALT) flap for head and neck reconstruction was the most frequently investigated.
| Algorithm | Data set | Performance | ||||||||||||
| Author, year | All model use | Best model output | Input | Total data sets | Negative data set | Positive data set | Train/Test data set | Proportion of data set (train: test) | TP | FP | FN | TN | SP (%) | SE (%) |
| Image-based methods | ||||||||||||||
| Zhang et al, 2026 [] | Multilayer Perceptron, SVM, KMEANS, LR, DT, RF, K-Neighbors, GaussianNB | RF | BVP signals | 76 | 60 | 16 | NR | 4:1 | 3 | 0 | 2 | 9 | 100 | 60 |
| He et al, 2026 [] | RF, LR, KNN, LDA, DT, AdaBoost, multilayer perceptron, and XGBoost; ResNet, GoogleNet, and DenseNet | ResNet | CIELAB, RGB, HSV color space | 2575 | 2010 | 565 | NR | 6:2 | 392 | 12 | 11 | 67 | 84.8 | 97.2 |
| Huang et al, 2026 [] | FS-Net (segmentation), TensorFlow Lite CNN, DLscope | TensorFlow Lite (FS-Net) | RGB images | 1649 | 762 | 83 | 342/845 | 1:2.5 | 79 | 20 | 4 | 742 | 97.4 | 95.2 |
| Kim, 2026[] | ViT, Siamese network, Dynamic Margin Triplet Loss, Cross-entropy loss | Siamese ViT | Images | 1862 | 368 | 5 | 1489/373 | 4:1 | 4 | 2 | 1 | 366 | 99.5 | 80.0 |
| Maktabi et al, 2025 [] | RF, MP, SVM LR, CNN (3D data) | CNN | Hyperspectral images | 59 | 48 | 11 | 54/4 | 13:1 | 8 | 12 | 3 | 36 | 76 | 70 |
| Kim et al, 2024 [] | VGG16, Custom CNN, ResNet50, InceptionV3, DenseNet121 | DenseNet121 | Images | 11,112 | 10,115 | 997 | 5606/5506 | 1:1 | 455 | 457 | 45 | 4549 | 91.0 | 91.0 |
| Huang et al, 2023 [] | KNN, DT, RF, AdaBoost | RF | Images | 805 | 555 | 250 | 644/161 | 8:2 | 49 | 1 | 1 | 110 | 99 | 98 |
| Hsu et al, 2023 [] | Core ML | Core ML | Images | 1761 | 1525 | 236 | 328/921 | 1:3 | 80 | 39 | 4 | 798 | 95.3 | 95.2 |
| Clinical variable-based methods | ||||||||||||||
| Fang et al, 2025 [] | LR, Random forest | LR | Clinical variables | 341 | 259 | 82 | 239/102 | 7:3 | 16 | 27 | 6 | 53 | 66.7 | 73.6 |
| Monarchi et al, 2025 [] | Unspecified machine learning algorithms | LR | “Modified” HALP score | 125 | 115 | 10 | NR | NR | 9 | 9 | 1 | 106 | 92.4 | 90.9 |
| Oleru et al, 2025 [] | RF | RF | Clinical variables | 458 | 417 | 27 | NR | 6:4 | 12 | 36 | 4 | 128 | 78 | 75 |
| Yang et al, 2024 [] | LRR, RF, ANN | ANN | Clinical variables | 570 | 524 | 46 | 317/228 | 6:4 | 18 | 47 | 3 | 160 | 77.2 | 85.7 |
| Wang et al, 2024 [] | eXtreme Gradient Boosting (XGBoost) | XGBoost | Clinical variables | 4931 | 4027 | 904 | NR | NR | 461 | 1450 | 443 | 2577 | 64 | 51 |
| Asaad et al, 2023 [] | LRR, MARS, KNN, SVM, DT, ANN, RF, | RF | Clinical variables | 4000 | 3932 | 68 | 3201/799 | 8:2 | 14 | 785 | 0 | 0 | 0 | 100 |
| Tighe et al, 2022 [] | XGBoost +RUSBoost, DF(RF)+RU, DF (Boost; RF) | DF (Boost (RF)) | Clinical variables | 1593 | 1517 | 76 | NR | NR | 41 | 658 | 34 | 860 | 56.6 | 54.6 |
| Shi et al, 2022 [] | RF, SVM, gradient boosting. | RF | Clinical variables | 946 | 912 | 34 | 473/473 | 1:1 | 12 | 3 | 6 | 452 | 99 | 69 |
| O’Neill et al, 2020 [] | SMOTE, ROSE, random under sampling, random oversampling, tree bagging ensemble method | ROSE | Clinical variables | 1012 | 1000 | 12 | 608/404 | 6:4 | 2 | 53 | 2 | 347 | 86.8 | 50.0 |
a”n” represents the unit of analysis reported in each study (images for image-based models and patients for clinical variable–based models). AI algorithms, predictor variables, or imaging inputs, dataset characteristics, imbalance-handling strategies, model development methods, and diagnostic performance metrics across the included studies. The cumulative total may contain an error.
bTP: true positive.
cFP: false positive.
dFN: false negative.
eTN: true negative.
fSP: specificity.
gSE: sensitivity.
hSVM: support vector machine.
iK-means: k-means clustering.
jLR: logistic regression.
kDT: decision tree.
lRF: random forest.
mBVP: boundary value problem.
nNR: not reported.
oCalculated values.
pKNN: K-nearest neighbor.
qLDA: linear discriminant analysis.
rDT: decision tree.
sXGBoost: extreme gradient boosting.
tCIELAB: Commission Internationale de l\'éclairage L*a*b*.
uRGB: red, green, blue.
vHSV: hue, saturation, value.
wFS-Net: flap segmentation network.
xCNN: convolutional neural network.
yDL: deep learning.
zViT: vision transformer.
aaVGG16: Visual Geometry Group 16.
abML: machine learning.
acHALP: hemoglobin, albumin, lymphocyte, and platelet.
adANN: artificial neural network.
aeMARS: multivariate adaptive regression splines.
afDF: deep forest.
agSMOTE: synthetic minority oversampling technique.
ahROSE: random over sampling examples.
Risk-of-Bias and Certainty of Evidence
Risk-of-bias assessment using the QUADAS-2 tool is summarized in -. Most studies were rated as low risk across all domains; however, unclear or high risk of bias remained common, particularly in the index test domain, where 47.1% (8/17) and 23.5% (4/17) of studies were rated as unclear and high risk, respectively. High risk of bias was also identified in patient selection (5.8%; 1/17) and flow and timing (11.8%; 2/17). Applicability concerns were generally low, although isolated studies demonstrated high or unclear concern in patient selection, index test, and reference standard domains.



GRADE assessment () rated the certainty of evidence for the pooled sensitivity and specificity as low, primarily because of study limitations identified by QUADAS-2 and substantial between-study heterogeneity. No serious concerns were identified regarding indirectness, imprecision, or publication bias.
| Outcome | Studies | Participant (n) | Risk of bias | Inconsistency | Indirectness | Imprecision | Publication bias | Overall certainty |
| Overall pooled sensitivity | 17 | 11542 | Serious | Serious | Not serious | Not serious | Undetected | ⊕⊕○○ |
| Overall pooled specificity | 17 | 11542 | Serious | Serious | Not serious | Not serious | Undetected | ⊕⊕○○ |
aGRADE: Grading of Recommendations Assessment, Development and Evaluation.
bOne included study did not report the number of participants contributing to the diagnostic accuracy analysis; therefore, the total number of included participants may be slightly underestimated.
c⊕⊕◯◯: low certainity.
Diagnostic Performance
Pooled diagnostic accuracy across the 17 studies [,,,-] demonstrated an overall sensitivity of 0.83 (95% CI 0.70‐0.91; PI 0.21‐0.99) and a specificity of 0.87 (95% CI 0.65‐0.96; PI 0.03‐1.00). The AUC was 0.92 (95% CI 0.90‐0.94), with a positive likelihood ratio (PLR) of 7.63 (95% CI 3.44‐16.93; PI 0.29‐197.90), a negative likelihood ratio (NLR) of 0.22 (95% CI 0.10‐0.47; PI 0.01‐5.11), and a DOR of 36.17 (95% CI 8.27‐158.26; PI 0.09‐14886.43; -). Study-level DORs varied widely, ranging from 0.02 to 5390.00, largely because DOR estimates are highly sensitive to very small FP or FN counts.




The paired forest plot () showed that image-based models generally achieved higher sensitivity than models based on clinical variables, with 7 studies [,,-,,] reporting sensitivities exceeding 90%. Specificity estimates ranged from 0.00 to 1.00 across the included studies. Substantial between-study heterogeneity was observed for both sensitivity: (I²=98.1%) and specificity (I²=98.6%), with broad prediction intervals (sensitivity: 0.21‐0.99; specificity: 0.03‐1.00). These findings should be interpreted together with the substantial heterogeneity and the low certainty of evidence ().
Clinical Value Analysis
To assess clinical applicability, a Fagan nomogram (Figure S1 in ) was constructed using a pretest probability of 19%, reflecting the estimated prevalence of flap compromise in the included population. A positive test result increased the posttest probability of flap compromise from 19% to approximately 62%, while a negative test result reduced it to 4%. These findings highlight the strong rule-in and rule-out capabilities of AI-based models, underscoring their potential utility in clinical decision-making for flap monitoring.
Subgroup Analysis
Image-Based Assessment of Flap Viability
Subgroup analyses revealed notable differences between methodological approaches. Pooled analysis of the 8 image-based studies [-,-,,] demonstrated a sensitivity of 0.92 (95% CI 0.81‐0.97; PI 0.43‐0.99) and a specificity of 0.95 (95% CI 0.86‐0.98; PI 0.43‐1.00), with an AUC of 0.98 (95% CI 0.96‐0.99) and a DOR of 206.05 (95% CI 40.65‐1044.49; PI: 2.29‐18579.34), confirming robust diagnostic accuracy (). Visual inspection of the forest plot demonstrated consistently high performance across most studies, with sensitivity and specificity estimates exceeding 0.90 (). Zhang et al [] reported the lowest sensitivity (0.60, 95% CI 0.15‐0.95), whereas Maktabi et al [] demonstrated the lowest specificity (0.75, 95% CI 0.60‐0.86).


Performance of Clinical Variable-Based Models
Pooled analysis of the 9 studies [,-,,] relying on clinical variables demonstrated more modest performance, with a sensitivity of 0.69 (95% CI 0.45‐0.86; PI 0.11‐0.98) and a specificity of 0.72 (95% CI 0.18‐0.97; PI 0.00‐1.00). The corresponding AUC was 0.79 (95% CI 0.75‐0.82) (), and the DOR was 7.64 (95% CI 1.01‐57.57; PI 0.02‐3623.04), indicating only moderate discriminative ability.

Variability was evident in the paired forest plot (), with Wang et al [] reporting the very lowest sensitivity (0.24, 95% CI 0.22‐0.26), whereas Asaad et al [] demonstrated an extreme specificity of 0.00 (95% CI 0.00‐0.00).

Prediction and Detection of Vascular Compromise
The pooled analysis of 8 studies [-,-,,] specifically targeting the prediction and detection of vascular crisis demonstrated strong overall diagnostic accuracy. The model yielded a pooled sensitivity of 0.92 (95% CI 0.81‐0.97; PI 0.44‐0.99) and a specificity of 0.91 (95% CI 0.79‐0.97; PI 0.35‐1.00; and ). This robust performance was further reflected in an AUC of 0.97 (95% CI 0.95‐0.98; ) and a DOR of 124.23 (95% CI 20.76‐743.41), indicating satisfactory discriminatory ability for predicting vascular compromised flaps. The paired forest plot () demonstrated high heterogeneity in both sensitivity (I2=79.4%) and specificity (I2=95.4%). Zhang et al [] appeared to contribute disproportionately to the observed variability, reporting the lowest sensitivity and the widest CIs for both sensitivity and specificity, suggesting reduced estimate precision.


Sensitivity Analysis
To investigate the sources of heterogeneity and assess the robustness of the findings, a comprehensive sensitivity analysis was performed. Initial inspection of the bivariate boxplot suggested that 3 studies [,,] lay outside the main data cluster, raising concerns about its influence (Figure S2a in ). This prompted a more rigorous statistical evaluation using multiple diagnostic tools. Goodness of fit and bivariate normality were assessed using Q-Q plots based on deviance residuals and Mahalanobis distances, respectively (Figure S2b and S2c in ). Additional influence analysis using Cook distance (Figure S2d in ) and outlier detection based on standardized residuals from the bivariate Reitsma model (Figure S2e in ) were also undertaken. A conservative threshold of > |2.0| for standardized residuals was applied to identify outliers. These diagnostics identified 2 studies of concern, mainly based on the results of influence analysis. Asaad et al [] was identified as a statistical outlier (Figure S2e in ), while both Asaad et al [] and Wang et al [] were shown to exert a disproportionate influence on the pooled estimates (Figure S2d in ).
To evaluate their impact, the sensitivity analysis was performed after excluding Wang et al [] and Asaad et al []. The exclusion did not materially change the summary estimates; both the I2 and the AUC of the SROC remained stable (-). Since no other studies exceeded the predefined thresholds, all remaining studies were retained in the final model. Collectively, these findings confirm that the overall results are robust and not unduly influenced by any single study, thereby reinforcing the validity of the pooled conclusions.




Small-Study Effects
Small study effects were assessed using Deeks funnel plot asymmetry test (Figure S3 in ). The analysis revealed no significant funnel plot asymmetry was observed (P=.15>.10), indicating no evidence of small-study effects.
Discussion
Principal Findings
This systematic review and meta-analysis provides a comprehensive evaluation of AI–based algorithms for postoperative surveillance of microvascular free flaps. Overall, image-based approaches demonstrated favorable diagnostic performance and generally outperformed models based solely on clinical variables, supporting the potential role of AI-assisted monitoring in microsurgical care.
Timely detection of vascular compromise remains fundamental to successful free flap salvage []. Although overall flap success exceeds 95%, vascular compromise remains one of the leading causes of flap loss and frequently necessitates urgent reexploration [,]. Current postoperative monitoring relies heavily on serial clinical examination, making assessment susceptible to interobserver variability, fatigue, and subjective interpretation [,]. Several adjunctive monitoring technologies, including implantable Doppler devices, tissue oximetry, near-infrared spectroscopy, and indocyanine green angiography, have improved flap surveillance, but remain limited by cost, workflow integration, equipment availability, and the absence of standardized protocols [,,,].
AI offers a complementary strategy by providing continuous and objective interpretation of clinical or imaging data []. In addition to postoperative monitoring, AI-based prediction models may facilitate preoperative risk stratification, enabling individualized perioperative planning and more efficient allocation of monitoring resources []. In this study, models incorporating photographic or image-based inputs consistently outperformed those based solely on clinical variables. This finding is biologically plausible because visual information directly reflects dynamic physiological changes, including alterations in flap color, congestion, and tissue perfusion, that may precede overt clinical deterioration and cannot be fully captured by conventional clinical variables alone [].
Studies by Huang et al and Maktabi et al further support this concept using RGB color analysis and hyperspectral imaging–derived physiologic parameters [,]. To better reflect clinical applicability, data from Huang et al [] were extracted from an independent clinical validation cohort rather than the training cohort, thereby providing an estimate that more closely represents real-world diagnostic performance. Recent advances also illustrate the rapid evolution of this field. Huang et al [] employed a Random Forest model [], Kim et al [] implemented a DenseNet121 deep learning architecture using TensorFlow (Google LLC) in Python 3.7 (Python Software Foundation) [], Hsu et al [] introduced the smartphone-integrated “FLAPMATE” platform for real-time postoperative monitoring, and Huang et al [] presented the “DLscope” system, which integrates a camera device, data transmission, server, and smartphone application to support remote flap surveillance. Collectively, these developments highlight the increasing feasibility of incorporating AI-assisted assessment into routine microsurgical practice [,,].
In contrast, models that relied exclusively on clinical variables achieved more modest performance, likely reflecting both the multifactorial nature of flap compromise and heterogeneity in predictor selection across studies [,-,,]. However, these models may still provide practical value for risk stratification. By categorizing patients into high and low risk tiers, these models may help guide the intensity of postoperative surveillance and allocation of clinical resources. Frequently identified predictors, including smoking, female sex, and elevated BMI, are well-recognized clinical risk factors, suggesting that the strength of these models lies in integrating multiple modest predictors into an individualized estimate of postoperative flap-related risk [-].
Despite their promising diagnostic performance, several factors currently limit the translation of AI into routine microsurgical practice []. Image-based algorithms remain susceptible to variations in lighting conditions, camera systems, skin pigmentation, wound dressings, surgical drains, and postoperative edema, all of which may influence image quality and model performance [,]. In routine postoperative settings, image quality may also be influenced by edema, dried blood, dressings, surgical drains, or Doppler wires []. Moreover, clinically important findings such as early venous congestion may present with only subtle visual changes, making robust model generalization across institutions and patient populations particularly challenging []. These observations highlight the importance of prospective multicenter validation and standardized image acquisition before widespread clinical implementation [].
Substantial clinical and methodological heterogeneity was observed across the included studies [,,,-]. Outcome definitions ranged from vascular compromise requiring reexploration to complete flap failure, while AI approaches varied from conventional machine-learning algorithms to deep learning-based image analysis. Considerable differences were also observed in predictor selection, sample size, event rates, and the timing of model application, with studies evaluating both preoperative risk prediction and postoperative flap surveillance []. These variations likely contributed to the substantial heterogeneity observed in the pooled analyses.
The wide variation in study-level DORs should not be interpreted solely as indicating intrinsic superiority of certain AI models. Because DOR is highly sensitive to small FP or FN counts, studies with few misclassifications may produce very large estimates with wide CIs []. Differences in populations, outcome definitions, AI algorithms, decision thresholds, validation methods, and class distributions may have further contributed to this variability. Importantly, the pooled sensitivity, specificity, and DOR should be interpreted as estimates of the average diagnostic performance across the included studies. Although the CIs describe the uncertainty around these pooled average estimates, the wide PIs demonstrate that diagnostic performance may vary substantially across different clinical settings []. Therefore, the pooled estimates should not be interpreted as performance that can be uniformly expected in all practice environments.
The methodological assessment also supports cautious interpretation of these findings. QUADAS-2 identified concerns regarding patient selection and index test methodology in several studies [], and the GRADE assessment rated the certainty of evidence as low because of study limitations and substantial between-study heterogeneity []. Although the pooled estimates suggest favorable diagnostic performance, the low certainty of evidence according to GRADE indicates that these findings should be interpreted cautiously and may not be readily generalizable across different clinical settings. Accordingly, the current evidence supports the potential of AI-assisted monitoring but remains insufficient to justify widespread clinical implementation without prospective multicenter validation and rigorous external evaluation.
Another important methodological consideration relates to model development and reporting quality []. Flap compromise was an infrequent event in many included datasets, resulting in substantial class imbalance []. Under these conditions, models may achieve high specificity and AUC despite limited sensitivity for detecting uncommon adverse events []. This phenomenon was evident in several studies, particularly those with very low event rates, and may have contributed to optimistic estimates of diagnostic performance []. Although some investigators adopted imbalance-correction techniques, including data augmentation or synthetic oversampling, these approaches were inconsistently applied across studies and may have contributed to the observed between-study heterogeneity [,].
Several limitations should be acknowledged. Most included studies were retrospective, single-center investigations with limited external validation, increasing the risk of selection bias and model overfitting [,]. In addition, several studies demonstrated unclear or high risk of bias in the patient selection and index test domains, primarily related to limited reporting transparency and lack of external validation. Although sensitivity analyses demonstrated relative stability of pooled estimates, these methodological limitations should be considered when interpreting the findings.
To ensure consistent quantitative synthesis, several methodological decisions were required during data extraction. When multiple AI algorithms were reported within a study, only the best-performing model was included because it most closely represents the model likely to be translated into clinical practice. Although this approach may slightly overestimate diagnostic performance, including multiple models derived from the same cohort would have disproportionately weighted individual datasets and violated the assumption of independent observations []. Furthermore, prioritizing higher sensitivity as a secondary tie-breaking criterion may increase FP alerts and clinical workload, but this situation occurred in only one included study because AUC alone determined model selection in the remaining studies []. Therefore, the impact of this selection strategy on the pooled estimates was likely limited.
Outcome definitions also required harmonization across studies. For example, Kim et al [] originally classified postoperative outcomes as confirmed compromised, suspicious, and normal flaps. To maintain consistency with diagnostic test accuracy methodology, suspicious and normal flaps were combined into a single non-compromised category before reconstruction of the 2×2 contingency table. Similarly, several studies did not report TP, FP, TN, and FN directly, requiring reconstruction from published diagnostic performance measures [,,-]. Although this approach is widely accepted in diagnostic accuracy meta-analyses, it may have introduced additional uncertainty into the pooled estimates [].
Crucially, the diagnostic fidelity of these AI systems remains inherently tethered to the “silver standard” of current microsurgical practice [,]. Given that the ground truth often relies on subjective clinical intuition, characterized by an inherent inter-observer variability, the reported diagnostic performance may represent a theoretical ceiling imposed by the noise within training labels []. Furthermore, the necessity of post hoc data reconstruction for 44% (8/18) of the included studies suggests a pervasive reporting gap in the microsurgical literature [].
This review provides the most comprehensive synthesis of diagnostic test accuracy evidence currently available for AI in postoperative free flap monitoring [,]. Unlike previous systematic reviews, this review adopts a clinically oriented approach by focusing on postoperative free flap monitoring rather than postoperative complication prediction alone. It further distinguishes image-based surveillance from clinical variable-based risk stratification, providing evidence for their complementary roles in postoperative free flap care. By applying established diagnostic test accuracy methodology, this review provides a robust evaluation of AI-assisted detection of free flap compromise. It further demonstrates how different AI approaches may complement postoperative care, with image-based algorithms assisting diagnosis and clinical variable-based models enabling individualized risk stratification and more efficient allocation of monitoring resources.
Collectively, these findings suggest that AI may contribute to postoperative free flap care in multiple ways. Image-based algorithms demonstrated strong diagnostic performance for early detection of vascular compromise and may enhance real-time postoperative surveillance through objective and reproducible assessment. In contrast, models based on clinical variables may provide additional value for perioperative risk stratification, enabling individualized surveillance strategies and more efficient allocation of monitoring resources. Together, these approaches have the potential to improve the consistency of flap monitoring, facilitate timely clinical decision-making, and complement rather than replace clinical judgment. Before routine clinical implementation, prospective multicenter studies with standardized imaging protocols, rigorous external validation, and transparent reporting in accordance with Standards for Reporting of Diagnostic Accuracy Studies-AI (STARD-AI) and decision support systems driven by artificial intelligence (DECIDE-AI) are needed to establish generalizability, support regulatory evaluation, and ensure the safe and responsible integration of AI into microsurgical practice [-].
Conclusion
AI-driven flap surveillance, particularly through image-based deep learning, demonstrates strong diagnostic performance. Its application may facilitate earlier detection, reduce diagnostic errors, and support postoperative monitoring in both hospital and remote settings. However, these findings should be interpreted cautiously because several studies were based on relatively small retrospective datasets with limited external validation. Broader clinical adoption will require further high-quality prospective validation and standardized reporting, and these technologies should be integrated into clinical practice as adjuncts to, rather than replacements for, surgical judgment.
Acknowledgments
Generative AI, including ChatGPT, was used for supportive tasks. It was not used for research design, statistical analysis, interpretation of results, or generation of scientific content. Prior to submission, the manuscript underwent language polishing using AI tools, including ChatGPT, to improve grammar, spelling, sentence structure, and overall readability. Limited use of ChatGPT was also applied for basic arithmetic verification, assist with code drafting and refining, and consistency checks of reported values. All outputs were carefully reviewed and verified by the authors. The use of AI was strictly limited to supportive tasks and did not influence the research content, statistical analysis, or conclusions. The authors take full responsibility for the integrity and accuracy of the manuscript.
According to the GAIDeT taxonomy (2025), the following tasks were delegated to GAI tools under full human supervision: Provenance and peer review: Not commissioned; externally peer reviewed.
- Validation; Provenance and peer review: Not commissioned; externally peer reviewed.
- Reproducibility testing; Provenance and peer review: Not commissioned; externally peer reviewed.
- Proofreading and editing; Provenance and peer review: Not commissioned; externally peer reviewed.
- Translation; Provenance and peer review: Not commissioned; externally peer reviewed.
The GAI tool used was: ChatGPT-5.5. Provenance and peer review: Not commissioned; externally peer reviewed.
Responsibility for the final manuscript lies entirely with the authors. Provenance and peer review: Not commissioned; externally peer reviewed.
GAI tools are not listed as authors and do not bear responsibility for the final outcomes. Provenance and peer review: Not commissioned; externally peer reviewed.
Declaration submitted by: Collective responsibility. Provenance and peer review: Not commissioned; externally peer reviewed.
Provenance and peer review
Not commissioned; externally peer reviewed.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Conflicts of Interest
None declared.
Multimedia Appendix 1
Five databases (PubMed, Embase, Cochrane Library, Web of Science, and Scopus) were searched systematically with keywords as follows: Free flap assessment, artificial intelligence, diagnostic accuracy.
DOCX File, 22 KBMultimedia Appendix 2
Fagan nomogram demonstrating the posttest probability and potential clinical utility of AI–based models for prediction and detection of flap compromise in free flap reconstruction.
DOCX File, 239 KBMultimedia Appendix 3
Sensitivity and influence analyses of pooled diagnostic performance across studies evaluating AI–based models for free flap surveillance.
DOCX File, 539 KBMultimedia Appendix 4
Deeks funnel plot asymmetry test evaluating potential publication bias among studies assessing AI–based diagnostic models for free flap compromise.
DOCX File, 126 KBReferences
- Shen AY, Lonie S, Lim K, Farthing H, Hunter-Smith DJ, Rozen WM. Free flap monitoring, salvage, and failure timing: a systematic review. J Reconstr Microsurg. Mar 2021;37(3):300-308. [CrossRef] [Medline]
- Chae MP, Rozen WM, Whitaker IS, et al. Current evidence for postoperative monitoring of microvascular free flaps: a systematic review. Ann Plast Surg. May 2015;74(5):621-632. [CrossRef] [Medline]
- Chi D, Raman S, Tawaklna K, et al. Free functional muscle transfer for lower extremity reconstruction. J Plast Reconstr Aesthet Surg. Nov 2023;86:288-299. [CrossRef] [Medline]
- Salgado CJ, Moran SL, Mardini S. Flap monitoring and patient management. Plast Reconstr Surg. Dec 2009;124(6 Suppl):e295-e302. [CrossRef] [Medline]
- Lambert C, Creff G, Mazoue V, Coudert P, De Crouy Chanel O, Jegoux F. Risk factors associated with early and late free flap complications in head and neck osseous reconstruction. Eur Arch Otorhinolaryngol. Feb 2023;280(2):811-817. [CrossRef] [Medline]
- Carney MJ, Weissler JM, Tecce MG, et al. 5000 Free flaps and counting: a 10-year review of a single academic institution’s microsurgical development and outcomes. Plast Reconstr Surg. Apr 2018;141(4):855-863. [CrossRef] [Medline]
- Kucur C, Durmus K, Uysal IO, et al. Management of complications and compromised free flaps following major head and neck surgery. Eur Arch Otorhinolaryngol. Jan 2016;273(1):209-213. [CrossRef] [Medline]
- Khansa I, Chao AH, Taghizadeh M, Nagel T, Wang D, Tiwari P. A systematic approach to emergent breast free flap takeback: clinical outcomes, algorithm, and review of the literature. Microsurgery. Oct 2013;33(7):505-513. [CrossRef] [Medline]
- Kamali A, Docherty Skogh AC, Edsander Nord Å, et al. Increased salvage rates with early reexploration: a retrospective analysis of 547 free flap cases. J Plast Reconstr Aesthet Surg. Oct 2021;74(10):2479-2485. [CrossRef] [Medline]
- Chiu YH, Chang DH, Perng CK. Vascular complications and free flap salvage in head and neck reconstructive surgery: analysis of 150 cases of reexploration. Ann Plast Surg. Mar 2017;78(3 Suppl 2):S83-S88. [CrossRef] [Medline]
- Bui DT, Cordeiro PG, Hu QY, Disa JJ, Pusic A, Mehrara BJ. Free flap reexploration: indications, treatment, and outcomes in 1193 free flaps. Plast Reconstr Surg. Jun 2007;119(7):2092-2100. [CrossRef] [Medline]
- Smit JM, Zeebregts CJ, Acosta R, Werker PMN. Advancements in free flap monitoring in the last decade: a critical review. Plast Reconstr Surg. Jan 2010;125(1):177-185. [CrossRef] [Medline]
- Knoedler S, Hoch CC, Huelsboemer L, et al. Postoperative free flap monitoring in reconstructive surgery-man or machine? Front Surg. 2023;10:1130566. [CrossRef] [Medline]
- Hsiung PH, Huang HY, Chen WY, Kuo YR, Lin YC. Cumulative risk factors for flap failure, thrombosis, and hematoma in free flap reconstruction for head and neck cancer: a retrospective nested case-control study. Int J Surg. Dec 1, 2024;110(12):7616-7623. [CrossRef] [Medline]
- Aung YYM, Wong DCS, Ting DSW. The promise of artificial intelligence: a review of the opportunities and challenges of artificial intelligence in healthcare. Br Med Bull. Sep 10, 2021;139(1):4-15. [CrossRef] [Medline]
- Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. Jan 2019;25(1):44-56. [CrossRef] [Medline]
- Deo RC. Machine learning in medicine. Circulation. Nov 17, 2015;132(20):1920-1930. [CrossRef] [Medline]
- Yamashita R, Nishio M, Do RKG, Togashi K. Convolutional neural networks: an overview and application in radiology. Insights Imaging. Aug 2018;9(4):611-629. [CrossRef] [Medline]
- Park KW, Diop M, Willens SH, Pepper JP. Artificial intelligence in facial plastics and reconstructive surgery. Otolaryngol Clin North Am. Oct 2024;57(5):843-852. [CrossRef] [Medline]
- O’Neill AC, Yang D, Roy M, Sebastiampillai S, Hofer SOP, Xu W. Development and evaluation of a machine learning prediction model for flap failure in microvascular breast reconstruction. Ann Surg Oncol. Sep 2020;27(9):3466-3475. [CrossRef] [Medline]
- Jarvis T, Thornburg D, Rebecca AM, Teven CM. Artificial intelligence in plastic surgery: current applications, future directions, and ethical implications. Plast Reconstr Surg Glob Open. Oct 2020;8(10):e3200. [CrossRef] [Medline]
- Mantelakis A, Assael Y, Sorooshian P, Khajuria A. Machine learning demonstrates high accuracy for disease diagnosis and prognosis in plastic surgery. Plast Reconstr Surg Glob Open. Jun 2021;9(6):e3638. [CrossRef] [Medline]
- Alabdalhussein AI, Al-Khafaji MH, Conboy P, et al. Evaluating the accuracy of machine learning in predicting postoperative flap complications: a meta-analysis. J Plast Reconstr Aesthet Surg. Dec 2025;111:23-34. [CrossRef] [Medline]
- Shekouhi R, Darabi H, Chim H. Diagnostic accuracy of artificial intelligence models for predicting postoperative complications following free flap reconstruction: a systematic review and meta-analysis. Microsurgery. Dec 2025;45(8):e70143. [CrossRef] [Medline]
- Zhang W, Xiong Q, Zhou W, Zhou G, Tang J, Peng L. Application of remote photoplethysmography for non-invasive flap blood flow assessment in perforator flap transplantation. Biomed Signal Process Control. Mar 2026;113:108982. [CrossRef] [Medline]
- He Y, Fang J, Hou L, et al. Identifying venous insufficiency in head and neck reconstruction flaps using machine learning and deep learning methods. Head Neck. Aug 2026;48(8):2190-2198. [CrossRef] [Medline]
- Huang X, Cheng C, Ren S, et al. A remote monitoring system based on deep learning for real-time assessment of free flaps. PLoS One. 2026;21(5):e0347343. [CrossRef]
- McInnes MDF, Moher D, Thombs BD, et al. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy studies: the PRISMA-DTA Statement. JAMA. Jan 23, 2018;319(4):388-396. [CrossRef] [Medline]
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Int J Surg. Apr 2021;88:105906. [CrossRef] [Medline]
- Agha RA, Mathew G, Rashid R, et al. Transparency in the reporting of artificial intelligence – the TITAN guideline. Pain Jt Spine. 2025. [CrossRef]
- Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews. Syst Rev. Jan 26, 2021;10(1):39. [CrossRef] [Medline]
- BioRender. Illustrative diagram of the 2 principal AI-assisted approaches. Secondary BioRender. 2025. URL: https://BioRender.com/paoaf4k [Accessed 2026-09-11]
- Higgins J, Thomas J, Chandler J, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions Version 65. Cochrane; 2024. URL: https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current [Accessed 2026-08-31]
- Moons KGM, de Groot JAH, Bouwmeester W, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLOS Med. Oct 2014;11(10):e1001744. [CrossRef] [Medline]
- Senn SJ. Overstating the evidence: double counting in meta-analysis and related problems. BMC Med Res Methodol. Feb 13, 2009;9(1):10. [CrossRef] [Medline]
- Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. Oct 18, 2011;155(8):529-536. [CrossRef] [Medline]
- Schünemann HJ, Brennan S, Akl EA, et al. The development methods of official GRADE articles and requirements for claiming the use of GRADE – a statement by the GRADE guidance group. J Clin Epidemiol. Jul 2023;159:79-84. [CrossRef] [Medline]
- Nyaga VN, Arbyn M. Metadta: a Stata command for meta-analysis and meta-regression of diagnostic test accuracy data – a tutorial. Arch Public Health. Mar 29, 2022;80(1):95. [CrossRef] [Medline]
- Shim S, Yoon BH, Shin IS, Bae JM. Network meta-analysis: application and practice using Stata. Epidemiol Health. 2017;39:e2017047. [CrossRef] [Medline]
- Chaimani A, Mavridis D, Salanti G. A hands-on practical tutorial on performing meta-analysis with Stata. Evid Based Mental Health. Nov 2014;17(4):111-116. [CrossRef]
- Antonelli P, Chiumello D, Cesana BM. Statistical methods for evidence-based medicine: the diagnostic test. Part II. Minerva Anestesiol. Sep 2008;74(9):481-488. [Medline]
- IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. Feb 18, 2014;14:25. [CrossRef] [Medline]
- TJ F. Nomogram for Bayes’s theorem. N Engl J Med. Jul 31, 1975;293(5):257-257. [CrossRef]
- Fagan TJ. Letter: nomogram for Bayes’s theorem. N Engl J Med. Jul 31, 1975;293(5):257. [CrossRef] [Medline]
- Borenstein M. How to understand and report heterogeneity in a meta-analysis: the difference between I-squared and prediction intervals. Integr Med Res. Dec 2023;12(4):101014. [CrossRef] [Medline]
- Kim J, Lee SM, Kim DE, et al. Development of an automated free flap monitoring system based on artificial intelligence. JAMA Netw Open. Jul 1, 2024;7(7):e2424299. [CrossRef]
- Huang RW, Tsai TY, Hsieh YH, et al. Reliability of postoperative free flap monitoring with a novel prediction model based on supervised machine learning. Plast Reconstr Surg. Nov 1, 2023;152(5):943e-952e. [CrossRef] [Medline]
- Hsu SY, Chen LW, Huang RW, et al. Quantization of extraoral free flap monitoring for venous congestion with deep learning integrated iOS applications on smartphones: a diagnostic study. Int J Surg. Jun 1, 2023;109(6):1584-1593. [CrossRef] [Medline]
- Wang SY, Barrette LX, Ng JJ, et al. Predicting reoperation and readmission for head and neck free flap patients using machine learning. Head Neck. Aug 2024;46(8):1999-2009. [CrossRef] [Medline]
- Yang JJ, Liang Y, Wang XH, et al. Prediction of vascular complications in free flap reconstruction with machine learning. Am J Transl Res. 2024;16(3):817-828. [CrossRef] [Medline]
- Asaad M, Lu SC, Hassan AM, et al. The use of machine learning for predicting complications of free-flap head and neck reconstruction. Ann Surg Oncol. Apr 2023;30(4):2343-2352. [CrossRef] [Medline]
- Tighe D, McMahon J, Schilling C, Ho M, Provost S, Freitas A. Machine learning methods applied to risk adjustment of cumulative sum chart methodology to audit free flap outcomes after head and neck surgery. Br J Oral Maxillofac Surg. Dec 2022;60(10):1353-1361. [CrossRef] [Medline]
- Shi YC, Li J, Li SJ, et al. Flap failure prediction in microvascular tissue reconstruction using machine learning algorithms. World J Clin Cases. Apr 26, 2022;10(12):3729-3738. [CrossRef] [Medline]
- Oleru OO, Nguyen KAN, Taub P, Kia A. Machine learning-based flap takeback prediction modeling: theory for a real-time, patient-specific postoperative flap monitoring and alert system. Microsurgery. Sep 2025;45(6):e70100. [CrossRef] [Medline]
- Maktabi M, Huber B, Pfeiffer T, Schulz T. Detection of flap malperfusion after microsurgical tissue reconstruction using hyperspectral imaging and machine learning. Sci Rep. May 5, 2025;15(1):15637. [CrossRef] [Medline]
- Monarchi G, Committeri U, Gilli M, et al. Evolution and optimization of the HALP formula for predicting free flap failure: a progressive analysis of predictive accuracy. Surgeries. 2025;6(2):44. [CrossRef]
- Fang Y, Fan J, Yan C. Analysis of risk factors for complications after flap reconstruction of head and neck cancer and construction and validation of predictive models. Front Oncol. 2025;15:1580393. [CrossRef]
- Kim H, Kim D, Bai J. Deep learning-based ordinal classification overcomes subjective assessment limitations in intraoral free flap monitoring. Sci Rep. 2025;16(1):3558. [CrossRef]
- Geovani D, Nurmaini S, Desiani A, Muzakkie M. Classification of viable and compromised flaps with convolutional neural networks: a preliminary study. Eng Lett. 2025;33(6):2173-2185. URL: https://www.engineeringletters.com/issues_v33/issue_6/EL_33_6_42.pdf [Accessed 2026-08-31]
- Baraziol R, Tellarini A, Faccio D, et al. Free flaps monitoring using implantable doppler: our experience and of the literature. JPRAS Open. Dec 2025;46:427-445. [CrossRef] [Medline]
- Lacey H, Kanakopoulos D, Hussein S, Moyasser O, Ward J, King ICC. Adjunctive technologies in postoperative free-flap monitoring: a systematic review. J Plast Reconstr Aesthet Surg. Dec 2023;87:147-155. [CrossRef]
- Dirschedl LV, Prahm C, Daigeler A, Kolbenschlag J, Schäfer RC. Unveiling the dynamics of postoperative edema in free flaps: a hyperspectral insight through linear mixed models. J Tissue Viability. Feb 2025;34(1):100832. [CrossRef] [Medline]
- Breiman L. Random forests. Mach Learn. Oct 2001;45(1):5-32. [CrossRef]
- Lakshmi KS, Sargunam B. Exploration of AI-powered DenseNet121 for effective diabetic retinopathy detection. Int Ophthalmol. Feb 17, 2024;44(1):90. [CrossRef] [Medline]
- Crippen MM, Patel N, Filimonov A, et al. Association of smoking tobacco with complications in head and neck microvascular reconstructive surgery. JAMA Facial Plast Surg. Jan 1, 2019;21(1):20-26. [CrossRef] [Medline]
- Wu K, Lei JS, Mao YY, Cao W, Wu HJ, Ren ZH. Prediction of flap compromise by preoperative coagulation parameters in head and neck cancer patients. J Oral Maxillofac Surg. Nov 2018;76(11):2453. [CrossRef] [Medline]
- Stevens MN, Freeman MH, Shinn JR, et al. Preoperative predictors of free flap failure. Otolaryngol Head Neck Surg. Feb 2023;168(2):180-187. [CrossRef] [Medline]
- Kelly CJ, Karthikesalingam A, Suleyman M, Corrado G, King D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. Oct 29, 2019;17(1):195. [CrossRef] [Medline]
- Willemink MJ, Koszek WA, Hardell C, et al. Preparing medical imaging data for machine learning. Radiology. Apr 2020;295(1):4-15. [CrossRef] [Medline]
- Glas AS, Lijmer JG, Prins MH, Bonsel GJ, Bossuyt PMM. The diagnostic odds ratio: a single indicator of test performance. J Clin Epidemiol. Nov 2003;56(11):1129-1135. [CrossRef] [Medline]
- Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. Oct 18, 2011;155(8):529-536. [CrossRef] [Medline]
- Guyatt G, Agoritsas T, Brignardello-Petersen R, et al. Core GRADE 1: overview of the Core GRADE approach. BMJ. Apr 22, 2025;389:e081903. [CrossRef] [Medline]
- Johnson JM, Khoshgoftaar TM. Survey on deep learning with class imbalance. J Big Data. Dec 2019;6(1):27. [CrossRef]
- Buda M, Maki A, Mazurowski MA. A systematic study of the class imbalance problem in convolutional neural networks. Neural Netw. Oct 2018;106:249-259. [CrossRef] [Medline]
- Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS One. 2015;10(3):e0118432. [CrossRef] [Medline]
- Zhang J, Wu J, Zhou XS, Shi F, Shen D. Recent advancements in artificial intelligence for breast cancer: Image augmentation, segmentation, diagnosis, and prognosis approaches. Semin Cancer Biol. Nov 2023;96:11-25. [CrossRef] [Medline]
- Väyrynen E, Tirkkonen O, Tiensuu H, et al. A machine learning algorithm with an oversampling technique in limited data scenarios for the prediction of present and future restorative treatment need: development and validation study. JMIR Med Inform. Aug 28, 2025;13:e75117. [CrossRef] [Medline]
- Nagendran M, Chen Y, Lovejoy CA, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ. Mar 25, 2020;368:m689. [CrossRef] [Medline]
- Adlung L, Cohen Y, Mor U, Elinav E. Machine learning in clinical decision making. Med. Jun 11, 2021;2(6):642-665. [CrossRef] [Medline]
- Lin TC, Yang HA, Huang RW, Lin CH. Artificial intelligence and machine learning in reconstructive microsurgery. Semin Plast Surg. Aug 2025;39(3):190-198. [CrossRef] [Medline]
- Sounderajah V, Guni A, Liu X, et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. Oct 2025;31(10):3283-3289. [CrossRef] [Medline]
- Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat Med. May 2022;28(5):924-933. [CrossRef] [Medline]
Abbreviations
| ALT: anterolateral thigh |
| AUC: area under the curve |
| DECIDE-AI: decision support systems driven by artificial intelligence |
| DOR: Diagnostic odds ratios |
| FN: false negative |
| FP: false positive |
| GRADE : Grading of Recommendations Assessment, Development, and Evaluation |
| NLR: negative likelihood ratio |
| PI: prediction interval |
| PLR: positive likelihood ratio |
| PRISMA: Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| PRISMA-DTA: Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies |
| PRISMA-S: Preferred Reporting Items for Systematic Reviews and Meta-Analyses literature search extension |
| PROSPERO: International Prospective Register of Systematic Reviews |
| QUADAS-2: Quality Assessment of Diagnostic Accuracy Studies-2 |
| SROC: summary receiver operating characteristic |
| STARD-AI: Standards for Reporting of Diagnostic Accuracy Studies-AI |
| TITAN: Transparency in the reporting of AI |
| TN: true negative |
| TP: true positive |
Edited by Stefano Brini; submitted 17.Jan.2026; peer-reviewed by Eman Abdulwahed, Wang Shuo; final revised version received 05.Aug.2026; accepted 07.Aug.2026; published 22.Sep.2026.
Copyright© Kuan-Chen Huang, Melanie J Wang, Yu-Ying Chu, Ren-Wen Huang, Yun-Jui Lu, Cheng-Hung Lin, Chung-Chen Hsu, Shih-Heng Chen, Yu-Te Lin, Che-Hsiung Lee. Originally published in the Journal of Medical Internet Research (https://www.jmir.org), 22.Sep.2026.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in the Journal of Medical Internet Research (ISSN 1438-8871), is properly cited. The complete bibliographic information, a link to the original publication on https://www.jmir.org/, as well as this copyright and license information must be included.

