The Turkish Language Association (TDK) is preparing a new dictionary expected to contain more than 800,000 words, TDK President Prof. Osman Mert told Anadolu Agency. The institution, based in the capital Ankara, is also developing a large language model for Turkish, a national term bank and three software tools that detect language errors.
Mert said preparing a dictionary of living Turkish is one of the main subjects TDK currently focuses on. "This will be a dictionary that represents all cultural fields and all branches of science, with a high capacity to represent Turkish," he said.
According to Mert, the current Turkish Dictionary contains about 132,000 words and covers the language of literature. He said initial trials for the new work indicate a vocabulary of more than 800,000 words.
A balanced corpus of one billion words compiled for the project will allow researchers to identify the meanings words acquire in context, as well as proverbs, idioms, terms and different functions of language structures, Mert said. TDK has 21 branches and commissions working on its projects.
Mert said TDK has worked for about three years on a large language model intended to train artificial intelligence that "thinks in Turkish". The theoretical stage is complete and the project has moved to the practical phase, he said.
The work is carried out with HAVELSAN. Mert said the technical infrastructure is largely finished and data entry has begun. The data is labeled in detail and the resulting dataset can also be used in language research, he added.
TDK first reactivated its terminology commission, Mert said, then formed an inter-institutional term coordination commission that includes ministries, public bodies and representatives of the Presidency.
The Türkiye Term Bank currently holds 310,000 terms and is planned to open this year to the public and relevant institutions, according to Mert. He said the software is designed to cover the Turkic world later, with terms from Turkic countries to be added to form a Turkic World Term Bank.
Mert said one program will identify language, spelling and punctuation errors in Turkish texts and suggest correct usage. A second program will scan public institutions' websites around the clock and report language errors to TDK.
A third program will identify words missing from the Turkish Dictionary. It will recognize all listed words, scan the internet throughout the day and report unlisted words to TDK. Mert said the three programs are planned to launch within one year.
Mert said TDK observes more careless use of spoken and written language among young people as the internet and social media spread. TDK cooperates with the Ministry of National Education on awareness activities.
TDK held diction courses for teachers first in Ankara, then in Cyprus, Karaman in south-central Türkiye and Denizli in the southwest, Mert said. TDK plans to extend the courses nationwide.
A story workshop held last year drew undergraduate and graduate students from various Ankara universities. The 120-hour program ran from October to June, and a selection of participants' works is now being printed. TDK issued a new call this year and has received many applications, Mert said. "Our aim is to support young people's writing skills in particular and to create awareness on this subject in society," he said.
Mert also addressed work on a common alphabet for the Turkic world. A 34-letter framework alphabet was adopted and announced in Baku, the capital of Azerbaijan, in 2024.
Academic boards of the countries concerned have evaluated the letters proposed for them since 2024, Mert said. The letters to be used by Kazakhstan and Kyrgyzstan were announced at a meeting in Astana, the capital of Kazakhstan, on June 15, 2026. The process in Uzbekistan continues.
"There is no change in the alphabet used by the Republic of Türkiye. There is no need for it either," Mert said.