Why SHAP and LIME Fall Short for Complex Models
SHAP and LIME remain the most widely used explainable AI methods, particularly for tabular data. Both rely on local approximations that, in their standard implementations, treat features as approximately independent and assume a linear decision boundary around the prediction point. Those assumptions break down for nonlinear, high-dimensional models. The linear approximation fails because the model's decision boundary is curved, and the independence assumption fails when features are correlated. In practice, these effects intensify with dimensionality. A perspective study in the biomedical domain found that both methods are highly affected by the choice of the underlying machine learning model and by feature collinearity. The study raised a note of caution on their usage and interpretation. Similar models can produce inconsistent explanations, which risks misleading downstream decisions. When features correlate, the attribution each method assigns can shift without any change in model behavior. The instability is not an edge case. It is a structural property of the approximation. Practitioners should not treat SHAP or LIME as universally reliable. The evidence points to selecting alternative or complementary interpretability tools based on model and data characteristics.
Concept-Based Explanations: TCAV and Beyond
Testing with Concept Activation Vectors (TCAV) interprets model decisions using human-level concepts rather than raw feature attributions. TCAV quantifies how strongly a concept influences a prediction by measuring the directional derivative of the model's output with respect to the concept's activation vector. To do this, you collect examples of the concept (e.g., images containing a certain object) and compare them to random examples. The method then tests whether the model's predictions shift systematically when that concept is present. This yields a global explanation: it tells you which concepts the model relies on across many inputs, not just one prediction.
TCAV is global: it tests concepts across many inputs. SHAP and LIME are primarily local, though SHAP can be aggregated. Attribution-based networks can provide both.
The trade-off is upfront work. Concepts must be defined a priori, which introduces selection bias. The choice of concepts shapes what the explanation can reveal. If you omit a relevant concept, the method will not detect it. This means the explanation is only as good as the concept list you provide. And the full scope of these trade-offs is still being characterized in the literature; the available studies note that these newer methods are still evolving.
Attribution-based interpretable neural networks offer another route. These networks are designed to produce explanations as part of their architecture, rather than as a post-hoc step. They provide both local and global explanations while maintaining classification performance comparable to black-box models. The constraint is architectural: these networks impose design requirements on the model itself, so you cannot apply them to an already-trained model without retraining or modifying it.
Concept-based methods suit bias audits and stakeholder communication, where human-level concepts matter more than token weights. For example, in a loan approval model, you can test whether the concept of "age" influences decisions by feeding the model activation vectors from examples of older and younger applicants. The answer you get is not which feature drove a prediction, but which concept the model relies on.
A Practical Framework for Choosing Interpretability Methods
Choosing an interpretability method starts with four factors: model architecture, task complexity, audience expertise, and regulatory requirements. Model architecture matters because some methods, like attribution-based networks, require a specific architecture. Task complexity covers dimensionality and nonlinearity; the earlier failure modes apply. Audience expertise determines the level of abstraction: regulators may need concept-level explanations, while engineers may prefer feature-level attributions. Regulatory requirements can also dictate the format, such as the need for counterfactual explanations in some jurisdictions.
For tabular data with high collinearity, avoid SHAP and LIME; the instability under collinearity, described earlier, makes them unreliable. Prefer concept-based methods or attribution-based interpretable networks instead. For high-dimensional, attention-based architectures, consider concept-based methods or attribution-based networks. Combine methods rather than relying on one. Use TCAV for bias audits, testing concepts like gender or race. Use attribution-based networks for both local and global explanations. Reserve simple feature attributions for low-risk, low-complexity models where the assumptions hold. For example, a credit risk team that needs to explain a single loan denial to a regulator would do better with an attribution-based network that gives local explanations, while a product team auditing for demographic bias would use TCAV to test concepts like race or gender.