Anar Kurmanzhankyzy recently introduced a collaborative initiative between the International Society “Qazaq Tili” and OpenAI. This project features the AI Evaluation Benchmark Suite, designed to assess large language models across various dimensions, such as comprehension, grammar, idiomatic expressions, and literary translation from Kazakh to English. Importantly, the benchmark is rooted in the Kazakh language’s unique linguistic and cultural attributes.
The project includes an ethnographic benchmark with 500 questions focused on Kazakh history and culture. The GPT-5.6 Terra model recorded a 35.92% accuracy rate in this evaluation.
Beyond evaluation, the initiative aims to safeguard Kazakhstan’s historical heritage by developing AI tools using optical character recognition (OCR) to digitize Kazakh-language archival content. This technology can process and extract text from various formats, including books and newspapers.
Additionally, developers are gathering Kazakh-language audio and text data for training purposes. Currently, there are 12,000 hours of cleaned audio recordings, and transcription models tailored for Kazakh achieve a remarkable 98% accuracy in speech-to-text conversions.
In previous reports, a 14-billion-token database was established to enhance ChatGPT’s Kazakh responses, highlighting the language’s rich historical and cultural context, including the Kazakh diaspora. Notably, Kazakh and Azerbaijani are among the fastest-growing languages in the ChatGPT platform, as noted by OpenAI’s CFO, Sarah Friar.
For further details, visit the Qazinform News Agency.