Uzbek Educational Text Corpus for TF–IDF-Based Text Analysis
Résumé fourni par la source
This dataset contains a corpus of 132 Uzbek-language educational texts collected and prepared for computational and statistical text analysis. The texts are stored in plain TXT format and organized into four archives: TXT_bolimlar1, TXT_bolimlar2, TXT_bolimlar3, and darslik. The corpus is intended to support research on Uzbek natural language processing (NLP), information retrieval, text mining, and quantitative analysis of educational texts. In particular, it can be used for experiments involving TF–IDF weighting, term–document matrix construction, vocabulary analysis, document similarity, keyword extraction, text classification, and related statistical and machine-learning methods. The dataset is particularly relevant to the study of Uzbek as an agglutinative language, where morphological variation can significantly affect vocabulary size, term frequency distributions, sparsity, and vector-space representations of documents. The files are provided in plain-text format to facilitate preprocessing and integration with common NLP and data-analysis tools. Researchers may apply their own tokenization, normalization, stop-word removal, stemming, lemmatization, or other preprocessing procedures depending on the objectives of their experiments. Dataset contents: TXT_bolimlar1.zip TXT_bolimlar2.zip TXT_bolimlar3.zip darslik.zip Total number of text documents: 132 The dataset was prepared for research on mathematical models and algorithms for intelligent analysis of Uzbek texts, with particular emphasis on TF–IDF-based methods.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.