Sayed, Imaan and Mahlaza, Zola and van der Leek, Alexander and Mopp, Jonathan and Keet, C. Maria (2025) On the usage of semantics, syntax, and morphology for noun classification in isiZulu, Proceedings of Third Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2025), 2 March 2025, Tallinn, Estonia, 96–105, University of Tartu Library, Estonia.
![]() |
Text
Resourceful2025.pdf Download (211kB) |
Abstract
There is limited work aimed at solving the core task of noun classification for Nguni languages. The task focuses on identifying the semantic categorisation of each noun and plays a crucial role in the ability to form semantically and morphologically valid sentences. The work by Byamugisha (2022) was the first to tackle the problem for a related, but non-Nguni, language. While there have been efforts to replicate it for a Nguni language, there has been no effort focused on comparing the technique used in the original work vs. contemporary neural methods or a number of traditional machine learning classification techniques that do not rely on human-guided knowledge to the same extent. We reproduce Byamugisha (2022)’s work with different configurations to account for differences in access to datasets and resources, compare the approach with a pre-trained transformer-based model, and traditional machine learning models that rely on less human-guided knowledge. The newly created data-driven models outperform the knowledge-infused models, with the best performing models achieving an F1 score of 0.97.
Item Type: | Conference paper |
---|---|
Subjects: | Computing methodologies > Artificial intelligence > Natural language processing Computing methodologies > Machine learning |
Date Deposited: | 13 Oct 2025 12:25 |
Last Modified: | 13 Oct 2025 12:25 |
URI: | https://pubs.cs.uct.ac.za/id/eprint/1754 |
Actions (login required)
![]() |
View Item |