Identifiant International des publications en série
et autres ressources périodiques, électroniques et imprimées

2019/09/18

Annif: DIY automated subject indexing using multiple algorithms

 Imprimer

 Télécharger

 Partager

 Envoyer à un(e) ami(e)

Manually indexing documents for subject-based access is a labour-intensive process. This paper describes Annif, an open source tool and microservice for automated subject indexing developed by the National Library of Finland. After training it with a subject vocabulary and existing metadata gathered from bibliographic databases, Annif can be used to assign subject headings for new documents. The current version is based on a combination of existing natural language processing and machine learning tools.