International Identifier for serials
and other continuing resources, in the electronic and print world

2019/09/18

Annif: DIY automated subject indexing using multiple algorithms

 Print

 Download

 Share

 Send to a friend

Manually indexing documents for subject-based access is a labour-intensive process. This paper describes Annif, an open source tool and microservice for automated subject indexing developed by the National Library of Finland. After training it with a subject vocabulary and existing metadata gathered from bibliographic databases, Annif can be used to assign subject headings for new documents. The current version is based on a combination of existing natural language processing and machine learning tools.