Identifying key contributors in open-source projects using social network analysis
Hakonen, Weeti (2025)
Diplomityö
Hakonen, Weeti
2025
School of Engineering Science, Tietotekniikka
Kaikki oikeudet pidätetään.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe20251214118744
https://urn.fi/URN:NBN:fi-fe20251214118744
Tiivistelmä
Social network analysis is a method that can be used to study collaboration patterns. Open-source software projects have numerous contributors, and these contributors often form relationships with one another through their contributions to the projects, ultimately forming a large social network. The goal of this thesis is to determine if social network analysis can be used to find the key contributors and characterize their roles in these networks.
The data for the study will be gathered from two large-scale open-source projects: Apache Spark and Microsoft Visual Studio Code. The interaction data is gathered from pull requests, and their affiliated comments and reviews from GitHub. The network will be analysed and formed based on this data.
The social network analysis was done using a weighted score for each user based on the contribution type. Centrality scores, specifically degree and betweenness centrality, were then calculated to identify the influence types for different contributors. The analysis successfully identified key contributors and aligned well with previous literature on how open-source projects tend to form. Sosiaalinen verkostoanalyysi on menetelmä, jota voidaan hyödyntää tutkimaan yhteistyömalleja. Avoimen lähdekoodin projekteissa on useita osallistuja, jotka muodostavat keskenään suhteita projektin sisällä tekemänsä työn kautta. Näiden suhteiden kautta syntyy sosiaalisia verkostoja. Tämän työn tavoitteena on selvittää, voidaanko sosiaalisen verkostoanalyysin avulla tunnistaa keskeisiä tekijöitä avoimen lähdekoodin projekteissa ja luonnehtia heidän roolejaan verkoston sisällä.
Tutkimuksen aineisto kerätään kahdesta avoimen lähdekoodin projekteista: Apache Spark ja Microsoft Visual Studio Code. Vuorovaikutusdata kerätään GitHubin pull requesteista, ja lisäksi niihin liittyvät kommentit kerätään myös. Itse verkosto muodostetaan ja analysoidaan näiden aineiston pohjalta.
Sosiaalinen verkostoanalyysi toteutettiin hyödyntämällä painotettua pistemäärää, joka määritettiin kontribuutiotyyppien perusteella. Keskeisyysarvot, kuten aste- ja välillisyyskeskeisyys laskettiin eri osallistujien vaikutustyyppien tunnistamiseksi. Analyysi onnistui tunnistamaan keskeiset osallistujat ja tulokset olivat yhteneväisiä aiemman kirjallisuuden kanssa siitä miten verkostot tyypillisesti muodostuvat.
The data for the study will be gathered from two large-scale open-source projects: Apache Spark and Microsoft Visual Studio Code. The interaction data is gathered from pull requests, and their affiliated comments and reviews from GitHub. The network will be analysed and formed based on this data.
The social network analysis was done using a weighted score for each user based on the contribution type. Centrality scores, specifically degree and betweenness centrality, were then calculated to identify the influence types for different contributors. The analysis successfully identified key contributors and aligned well with previous literature on how open-source projects tend to form.
Tutkimuksen aineisto kerätään kahdesta avoimen lähdekoodin projekteista: Apache Spark ja Microsoft Visual Studio Code. Vuorovaikutusdata kerätään GitHubin pull requesteista, ja lisäksi niihin liittyvät kommentit kerätään myös. Itse verkosto muodostetaan ja analysoidaan näiden aineiston pohjalta.
Sosiaalinen verkostoanalyysi toteutettiin hyödyntämällä painotettua pistemäärää, joka määritettiin kontribuutiotyyppien perusteella. Keskeisyysarvot, kuten aste- ja välillisyyskeskeisyys laskettiin eri osallistujien vaikutustyyppien tunnistamiseksi. Analyysi onnistui tunnistamaan keskeiset osallistujat ja tulokset olivat yhteneväisiä aiemman kirjallisuuden kanssa siitä miten verkostot tyypillisesti muodostuvat.
