FROM AN ANNOTATED CORPUS TO AN ONTOLOGY: A FORMAL MODEL OF MULTI-LEVEL SEMANTIC ANNOTATION OF FAKE NEWS IN THE KAZAKH–RUSSIAN MEDIA SPACE (KazFakeCorpus / KazFakeOnto)

Main Article Content

Anargul Ajmuratovna Nekessova
Elena Anatolevna Sidorova
Zhanar Bejbutovna Lamasheva
Madina Aralbaevna Sambetbayeva
Aigerim Sembekovna Yarimbetova
Mira ZHorabekovna Kaldarova
Aksaule Abzalқyzy Nazymkhan

Abstract

The paper addresses the formal representation of knowledge about disinformation in a low-resource bilingual media environment. Automatic fake news detection is traditionally formulated as a binary REAL/FAKE classification task, which does not capture the internal structure of an unreliable message: the type of fake content, the disinformation technique employed, the author’s communicative intent, and the characteristics of the source and evidence base.


Two interrelated resources are presented. The first is KazFakeCorpus, a bilingual corpus of 4,276 texts in Kazakh and Russian, balanced by class within each language; the REAL class was drawn from official materials of the Gov.kz portal and the FAKE class from their controlled transformations. The texts were annotated in Label Studio by two independent experts following a multi-level semantic scheme; Krippendorff’s Alpha for the principal levels ranged from 0.79 to 0.88. The second resource is KazFakeOnto, an ontology that recasts the annotation scheme from a flat relational representation into a formal OWL 2 DL model: 38 classes, 26 object properties, 22 data properties, 112 individuals and 1,155 asserted triples. Eleven defined classes are not asserted manually but inferred by a reasoner from property values; a property chain links a news item to its annotator through an annotation object. Consistency was verified with HermiT, and the query layer was evaluated on fifteen SPARQL and thirteen DL queries.


The ontological layer is shown to provide a reproducible specification of the annotation scheme, the automatic inference of semantic categories, and a unified formal environment for formulating corpus-linguistic questions and analysing machine-learning model errors. The limitations of the approach and directions for scaling the ontology to the full corpus and to naturally occurring disinformation are discussed.

Article Details

How to Cite
Nekessova, A. A., E. A. Sidorova, Z. B. Lamasheva, M. A. Sambetbayeva, A. S. Yarimbetova, M. Z. Kaldarova, and A. A. Nazymkhan. “FROM AN ANNOTATED CORPUS TO AN ONTOLOGY: A FORMAL MODEL OF MULTI-LEVEL SEMANTIC ANNOTATION OF FAKE NEWS IN THE KAZAKH–RUSSIAN MEDIA SPACE (KazFakeCorpus / KazFakeOnto)”. Russian Digital Libraries Journal, vol. 29, no. 6, Oct. 2026, pp. 2158-87, doi:10.26907/1562-5419-2026-29-6-2158-2187.


Most read articles by the same author(s)