A Comparative Explainable Machine Learning Framework for Predicting QRIS Adoption Using Synthetic MSME Profiles
- Publication History
- Published online: October 31, 2026
- DOI
- https://doi.org/10.35877/454RI.daengku5252
- Copyright
- Copyright (c) 2026 Erni Febrina Harahap, S Salfadri, Jhon Rinaldo
- User License
- https://creativecommons.org/licenses/by-nc-sa/4.0
Abstract
Quick Response Code Indonesian Standard (QRIS) has become the national interoperable payment infrastructure for Indonesian micro, small, and medium enterprises (MSMEs), yet detailed public microdata linking MSME characteristics to adoption status remain unavailable, which limits the development of predictive analytics for merchant onboarding. This study develops and evaluates a comparative explainable machine learning framework for predicting potential QRIS adoption using theory-informed synthetic MSME profiles generated through Monte Carlo simulation. A transparent data-generating process combining linear, quadratic, interaction, and threshold components produced 2,000 profiles under a moderate digital environment and 2,000 additional profiles under digitally enabling and digitally constrained scenarios. Four algorithmic families were compared: Logistic Regression, Support Vector Machine with radial basis function kernel (SVM-RBF), CatBoost, and the Explainable Boosting Machine (EBM). Evaluation integrated discrimination, probability calibration, distribution-shift testing, permutation feature importance, Accumulated Local Effects, intrinsic EBM feature functions, counterfactual explanations, and robustness analysis over sample size, label noise, class imbalance, coefficient perturbation, and random seeds. On the independent test set, SVM-RBF achieved the highest balanced accuracy (0.767) and the lowest Brier score (0.173), while Logistic Regression attained the highest ROC-AUC (0.821); the intrinsically interpretable EBM remained within approximately one percentage point of the best model on every metric. Feature-importance rankings were highly consistent across the four model families (Spearman ? ? 0.91), with facilitating conditions, digital financial literacy, and performance expectancy dominating the simulated predictions. Because all profiles are synthetic, the reported probabilities and feature effects are computational outcomes under explicitly stated simulation assumptions rather than empirical estimates or causal evidence. The framework demonstrates that a transparent, calibrated, and robustness-audited prediction pipeline for QRIS adoption is computationally feasible and provides a reproducible template for future validation with real MSME data.
Keywords
Citation
Statements

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.