AI-Based Metric for the Scientific Text Novelty
Kruzenshtern, Anna; Vuktor, Dodonov; Chechurin, Leonid (2025-10-29)
Huom!
Sisältö avataan julkiseksi: 30.10.2026
Sisältö avataan julkiseksi: 30.10.2026
Post-print / Final draft
Kruzenshtern, Anna
Vuktor, Dodonov
Chechurin, Leonid
29.10.2025
775
367-378
Springer, Cham
IFIP Advances in Information and Communication Technology
School of Engineering Science
Kaikki oikeudet pidätetään.
© 2026 IFIP International Federation for Information Processing
© 2026 IFIP International Federation for Information Processing
Julkaisun pysyvä osoite on
https://urn.fi/URN:NBN:fi-fe20251209116145
https://urn.fi/URN:NBN:fi-fe20251209116145
Tiivistelmä
The study introduces an information-theoretic approach for the quantitative assessment of unpredictability and, by extension, the originality of ideas. Two metrics are proposed, both of which are potentially useful for analyzing scientific or engineering texts.
The first metric, Semantic Gain, quantifies the amount of new knowledge contained in a text relative to the knowledge embedded in a baseline model. The second metric captures informational unexpectedness (or surprisal score), which is grounded in information entropy and measures how surprising a given text is to a well-trained language model.
Both metrics are applied to the abstracts of peer-reviewed engineering articles, and the results demonstrate that, despite being derived from different computational approaches and model architectures, the two metrics exhibit a strong linear correlation. This correlation may be interpreted as evidence of their mutual consistency and conceptual validity.
The first metric, Semantic Gain, quantifies the amount of new knowledge contained in a text relative to the knowledge embedded in a baseline model. The second metric captures informational unexpectedness (or surprisal score), which is grounded in information entropy and measures how surprising a given text is to a well-trained language model.
Both metrics are applied to the abstracts of peer-reviewed engineering articles, and the results demonstrate that, despite being derived from different computational approaches and model architectures, the two metrics exhibit a strong linear correlation. This correlation may be interpreted as evidence of their mutual consistency and conceptual validity.
Lähdeviite
Kruzenshtern, A., Dodonov, V., Chechurin, L. (2026). AI-Based Metric for the Scientific Text Novelty. In: Cavallucci, D., Brad, S., Livotov, P., Houssin, R. (eds) World Conference of AI-Powered Innovation and TRIZ Methodology. TFC 2025. IFIP Advances in Information and Communication Technology, vol 775. Springer, Cham. https://doi.org/10.1007/978-3-032-08851-2_24
Alkuperäinen verkko-osoite
https://link.springer.com/chapter/10.1007/978-3-032-08851-2_24Kokoelmat
- Tieteelliset julkaisut [1857]