Hyppää sisältöön
    • Suomeksi
    • På svenska
    • In English
  • Suomeksi
  • In English
  • Kirjaudu
Näytä aineisto 
  •   Etusivu
  • LUTPub
  • Diplomityöt ja Pro gradu -tutkielmat
  • Näytä aineisto
  •   Etusivu
  • LUTPub
  • Diplomityöt ja Pro gradu -tutkielmat
  • Näytä aineisto
JavaScript is disabled for your browser. Some features of this site may not work without it.

Assessing the trustworthiness and acceptability of explanations for machine learning

Aziz, Umair (2026)

Katso/Avaa
Mastersthesis_Aziz_Umair.pdf (2.099Mb)
Lataukset: 


Diplomityö

Aziz, Umair
2026

School of Engineering Science, Tietotekniikka

Näytä kaikki kuvailutiedot
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe20260618100392

Tiivistelmä

As machine learning systems become increasingly complex, Explainable Artificial Intelligence (XAI) has emerged to improve the transparency and accountability of model behavior. A central challenge in this field is the persistent tension between two competing qualities, faithfulness (the extent to which an explanation accurately reflects the internal decision making of the model), and plausibility (the extent to which an explanation appears coherent and convincing to human users). This tension has become more pronounced with the development of Large Language Models (LLMs) trained using reinforcement learning from human feedback (RLHF), as these systems are optimized to produce responses that align with human preferences, and it often results in fluent but potentially misleading explanations that may convey a sense of credibility without accurately representing the underlying computation.

This thesis empirically examines whether human users can reliably distinguish between faithful and only plausible AI-generated explanations and how different explanation strategies influence perceived trustworthiness in the context of automated hate speech detection. An experimental design within the subject was conducted in which eight participants evaluated 100 hate speech classification instances under four explanation conditions, a BERT-only classification output, a standalone GPT-generated explanation, a GPT explanation conditioned on the BERT label, and a GPT explanation enhanced with SHAP-based token importance scores. Participants rated each explanation on three five-point Likert scales measuring plausibility, faithfulness, and trustworthiness, resulting in a total of 9,600 individual ratings. The study used three datasets, HateXplain, Trawling for Trolling, and Hatebase with HateXplain providing token-level rationales used to define ground-truth faithfulness.

The results reveal a consistent pattern of over-trust, all GPT-based explanation methods were rated significantly higher than the BERT baseline across all three dimensions (all p < .001, Cohen’s d = 0.48–0.67). However, trust ratings barely changed depending on whether the underlying classification was correct, trust under incorrect classifications stayed within 96– 98% of trust under correct ones, showing that participants did not lower their trust when the system was wrong. Importantly, participants were unable to reliably distinguish between faithful and unfaithful explanations, as the correlations between ground-truth faithfulness and human-rated faithfulness were close to zero (all r < 0.03, p > 0.43). While SHAP augmentation led to a small but statistically significant increase in perceived faithfulness (Δ ≈ +0.17), it did not improve participants’ ability to identify whether an explanation was actually aligned with the ground-truth rationale tokens.

Contrary to expectations, GPT explanations did not systematically prioritize the plausibility over perceived faithfulness. Instead, the GPT+SHAP condition uniquely reversed this trend. Overall, the findings suggest that linguistic fluency strongly influences perceived credibility, even when explanations are not faithful to the underlying model. The study provides a structured four-condition framework for evaluating XAI systems and highlights the risks of deploying generative explanation methods in high-stakes domains without explicit mechanisms to verify faithfulness. These results have important implications for the design of accountable AI systems and the governance of automated content moderation.
Kokoelmat
  • Diplomityöt ja Pro gradu -tutkielmat [15472]
LUT-yliopisto
PL 20
53851 Lappeenranta
Ota yhteyttä | Tietosuoja | Saavutettavuusseloste
 

 

Tämä kokoelma

JulkaisuajatTekijätNimekkeetKoulutusohjelmaAvainsanatSyöttöajatYhteisöt ja kokoelmat

Omat tiedot

Kirjaudu sisäänRekisteröidy
LUT-yliopisto
PL 20
53851 Lappeenranta
Ota yhteyttä | Tietosuoja | Saavutettavuusseloste