Propensity Score Matching for Causal Inference: Comparing Classical and Bayesian Approaches

Naomi Kollongei *

Department of Mathematics and Computer Science, University of Eldoret, Eldoret, Kenya.

Argwings Otieno

Department of Mathematics and Computer Science, University of Eldoret, Eldoret, Kenya.

Julius Koech

Department of Mathematics and Computer Science, University of Eldoret, Eldoret, Kenya.

*Author to whom correspondence should be addressed.


Abstract

Propensity-score matching (PSM) is widely used to reduce measured confounding in observational studies, but treatment-effect estimates can be sensitive to matching restrictions and to uncertainty in the estimated propensity score. This study compares classical and Bayesian PSM using the 996-record Lindner abciximab dataset distributed with the LocalControl R package. The treatment was abciximab use during percutaneous coronary intervention, and the outcome was life-years preserved (lifepres), coded 0 for patients who died within 1 year and 11.6 for patients who survived at least 1 year in this public data distribution. The reported analysis used a six-covariate treatment-assignment model (stent use, height, sex, diabetes, recent acute myocardial infarction, and left ventricular ejection fraction). Classical analyses used 1:1 nearest-neighbour matching without replacement, first unrestricted and then with a caliper of 0.2 standard deviations of the logit propensity score. The classical model had an AUC of 0.6192 (95% CI 0.5820-0.6565). Unrestricted matching paired 298 of 698 treated patients with 298 controls and yielded a matched-sample mean difference of 0.3503 life-year units but poor residual balance. Caliper matching retained 288 treated-control pairs and yielded 0.4833 life-year units; balance improved materially for several covariates, although residual imbalance remained for stent use, height, sex, and the propensity score. Bayesian caliper matching propagated uncertainty in the treatment-assignment model through repeated posterior propensity-score draws and matching, generating 1,000 matched-sample treatment-effect draws. The posterior mean was 0.4098 life-year units (SD 0.2340; 95% credible interval -0.0410 to 0.8736), with P(effect > 0 | data) = 0.960. These results illustrate that matching restrictions can alter both the analysed treated population and the estimated treatment contrast. Given the remaining important residual imbalance, the omission of the number of vessels involved in the first PCI from the fitted treatment model, and the possibility of unmeasured confounding, the empirical estimates should be interpreted as a methodological illustration rather than as definitive clinical causal effects.

Keywords: Propensity score matching, Bayesian propensity score, nearest-neighbour matching, caliper matching, matched-sample treatment effect, causal inference, observational data


How to Cite

Kollongei, Naomi, Argwings Otieno, and Julius Koech. 2026. “Propensity Score Matching for Causal Inference: Comparing Classical and Bayesian Approaches”. Journal of Advances in Mathematics and Computer Science 41 (10):1-8. https://doi.org/10.9734/jamcs/2026/v41i102205.

Downloads

Download data is not yet available.