Skip to main navigation Skip to search Skip to main content

Joint distribution-informed shapely values for sparse counterfactual explanations

  • University of Copenhagen
  • Microsoft Gaming

Research output: Chapter in Book/Report/Conference proceedingArticle in proceedingsResearchpeer-review

Abstract

Counterfactual explanations (CE) aim to reveal how small input changes flip a model’s prediction, yet many methods modify more features than necessary, reducing clarity and actionability. We introduce COLA, a model- and generator-agnostic post-hoc framework that refines any given CE by computing a coupling via optimal transport (OT) between factual and counterfactual sets and using it to drive a Shapley-based attribution (p-SHAP) that selects a minimal set of edits while preserving the target effect. Theoretically, OT minimizes an upper bound on the W1 divergence between factual and counterfactual outcomes and that, under mild conditions, refined counterfactuals are guaranteed not to move farther from the factuals than the originals. Empirically, across four datasets, twelve models, and five CE generators, COLA achieves the same target effects with only 26–45% of the original feature edits. On a small-scale benchmark, COLA shows near-optimality
Original languageEnglish
Title of host publicationProceedings of 14th International Conference on Learning Representations
Number of pages23
Publication statusAccepted/In press - 2026
Event14th International Conference on Learning Representations - Rio de Janeiro, Brazil
Duration: 23 Apr 202627 Apr 2026

Conference

Conference14th International Conference on Learning Representations
Country/TerritoryBrazil
CityRio de Janeiro
Period23/04/202627/04/2026

Fingerprint

Dive into the research topics of 'Joint distribution-informed shapely values for sparse counterfactual explanations'. Together they form a unique fingerprint.

Cite this