Original Article

TOLID: Turkish Offensive Language Identification Dataset and Transformer-Based Benchmarks

Volume 26 Publish Date: July 24, 2026
Full Text PDF
DOI
Mehmet Salih Kurt ORCID
Department of Computer Engineering, Hakkari University Faculty of Engineering, Hakkari, Türkiye
Eylem Yücel ORCID
Department of Computer Engineering, İstanbul University-Cerrahpaşa Faculty of Engineering, İstanbul, Türkiye
Kurt, M. S., & Yücel, E. (2026). TOLID: Turkish Offensive Language Identification Dataset and Transformer-Based Benchmarks. ELECTRICA, 26, 1–10. https://doi.org/10.5152/electrica.2026.25404
Full Text Full-Text PDF

Abstract

This study presents the Turkish Offensive Language Identification Dataset (TOLID), a large-scale, high-quality dataset for the automatic detection of offensive language in Turkish social media posts and evaluates the performance of several transformer-based models. Although most previous studies have been constrained by small sample sizes, imbalanced class distributions, or narrow topical focus, TOLID includes a wide range of offensive expressions without imposing restrictions on topics, individuals, or groups. The dataset was annotated by three independent experts using a hierarchical and fine-grained scheme that addresses not only general offensive language but also specific subcategories such as sexist, racist, political, and religious insults. To evaluate the dataset, several transformer-based models adapted for Turkish, including BERTurk (Bidirectional Encoder Representations from Transformers for Turkish), ConvBERTurk (Convolutional Bidirectional Encoder Representations from Transformers for Turkish), and ELECTRA-Turkish (Efficiently Learning an Encoder that Classifies Token Replacements Accurately for Turkish), were trained and tested. Among these, ConvBERTurk achieved the highest scores, reaching 82.76% macro F1 in offensive vs. non-offensive classification and 78.14% in targeted vs. non-targeted offensive classification, outperforming previous research. These results demonstrate that combining a balanced, multi-annotated dataset with advanced deep learning architectures can effectively address the linguistic richness, contextual complexity, and informal nature of Turkish social media text. Additionally, a web-based application was developed to provide a practical interface for analyzing text and visualizing model outputs, extending the study's impact beyond academic research to real-world applications. Overall, this study addresses key limitations of prior research and makes a significant contribution to Turkish natural language processing by providing a meticulously constructed dataset and extensive benchmarks with state-of-the-art models. TOLID establishes a robust foundation for future work on offensive language subtypes, automatic moderation, and toxicity analysis in Turkish social media.

Cite this article as: M. S. Kurt and E. Yücel, “TOLID: Turkish offensive language identification dataset and transformer-based benchmarks,” Electrica, 26, 0404, 2026. doi: 10.5152/electrica.2026.25404.

 

Article Info
Published In
Journal ELECTRICA
Volume / Issue Volume 26
Pages 1-10
History
Published Online July 24, 2026
Affiliations
Mehmet Salih Kurt ORCID
Department of Computer Engineering, Hakkari University Faculty of Engineering, Hakkari, Türkiye
Eylem Yücel ORCID
Department of Computer Engineering, İstanbul University-Cerrahpaşa Faculty of Engineering, İstanbul, Türkiye
Cite this Article
Kurt, M. S., & Yücel, E. (2026). TOLID: Turkish Offensive Language Identification Dataset and Transformer-Based Benchmarks. ELECTRICA, 26, 1–10. https://doi.org/10.5152/electrica.2026.25404
Share
Outlines