Fine-grained propaganda analysis : span detection and technique classification
Ahmed, Omer (2026)
Diplomityö
Ahmed, Omer
2026
School of Engineering Science, Laskennallinen tekniikka
Kaikki oikeudet pidätetään.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe2026061268858
https://urn.fi/URN:NBN:fi-fe2026061268858
Tiivistelmä
This thesis investigates automatic propaganda detection in news articles written in English, addressing two subtasks: span identification and technique classification. The rapid growth of online information platforms has made large-scale propaganda analysis increasingly important. Manual detection is no longer an option, considering how fast propaganda spreads in the digital world. This study examines how propaganda detection methods have evolved over time and explores how existing techniques can be combined to improve performance. For span identification, seven Robustly Optimized BERT Pretraining Approach (RoBERTa)-large models are fine-tuned and evaluated, incorporating progressively richer feature sets including part-of-speech tags, named entity labels, and discourse-level features, combined with a Conditional Random Field (CRF) decoder. The best-performing model achieved an F1-score of 0.4817 on the SemEval-2020 Task 11 development set which is competitive with earlier published approaches but below the top systems. For technique classification, a span-aware pooling strategy with inverse-frequency class weighting was proposed, achieving a micro-averaged F1-score of 0.8878 across 14 propaganda techniques. A few-shot Generative Pre-trained Transformer (GPT)-based baseline was also evaluated on both tasks, achieving competitive recall but substantially lower precision than the fine-tuned models. Embedding analysis reveals that discourse and sequential features produce more separable latent representations, and identifies a cluster of implicitly framed propaganda spans that remain difficult for all models to detect. All the results demonstrate that fine-tuned transformer-based models outperform models (LLMs) for precise propaganda detection, while linguistic and discourse features provide meaningful complementary signals beyond contextual embeddings alone.
