Exploring streaming video understanding for figure skating
Ali, Haider (2026)
Diplomityö
Ali, Haider
2026
School of Engineering Science, Laskennallinen tekniikka
Kaikki oikeudet pidätetään.
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe2026060865039
https://urn.fi/URN:NBN:fi-fe2026060865039
Tiivistelmä
While Video Large Language Models (Video-LLM) perform well in offline video understanding, their use in streaming video, where frames arrive continuously and outputs must be produced on the fly, is still a growing field. This thesis investigates how a streaming Video-LLM behaves in figure skating by adapting the VideoLLM-Online framework and its LIVE (Learning-In-Video-Stream) method to element recognition from broadcast video. Using the FSAnno dataset of 7,221 element clips from 11 International Skating Union (ISU) Grand Prix competitions, the model Meta-Llama-3-8B-Instruct with Low Rank Adaptation (LoRA) is studied both offline, across vision encoders, temporal modeling, task decomposition and training signals, and in a streaming setting. Coarse grained understanding (Jump, Spin, Sequence) was effectively solved at 94.78% accuracy, subcategory classification of 21 types reached 54.75%, and the full 269 element codes reached at best 28.12% with the dual encoder SigLIP 2 + DINOv2. In streaming, the model decides when to respond, stays silent between elements, self-corrects as frames arrive, and times its responses to element boundaries, though segmenting a full performance remains open. The main bottleneck is that vision encoders process each frame independently and miss the motion detail that separates similar elements. A web-based annotation tool was also built to extend the dataset with the 2024-25 ISU Grand Prix data.
