Are AI Detection Tools Reliable on Turkish Texts?

Are AI Detection Tools Reliable on Turkish Texts?

You walk into the classroom and a piece of homework comes in from a pupil. The sentences are polished, the transitions clean, the word choices a touch too “mature”. A question forms in your mind: did ChatGPT write this?

Our first reflex is to paste the text into an AI detection tool (an AI detector) and wait for the verdict. Turnitin, GPTZero, ZeroGPT, Originality.ai, Copyleaks. The page says “87% written by AI”. Case closed?

No. Because every one of these tools has a serious problem with Turkish.

How do AI detection tools work?

AI detectors look at the statistical patterns in a text. Two core measures are used:

Perplexity: This measures how “unexpected” a text is. Text generated by AI generally makes more predictable word choices, which produces a low perplexity score.

Burstiness: In human writing, long and short sentences fluctuate irregularly. AI tends to produce smoother, more uniform sentence lengths.

These two measures give reasonable results for English, because detection tools were overwhelmingly trained on English data sets. For Turkish, the picture is different.

Why are they unreliable on Turkish?

There are three fundamental problems.

1. The language bias problem

According to a study published in 2023 by researchers at Stanford University, AI detectors wrongly flag texts by non-native English speakers as “AI-generated” at an average rate of 61.3%. For native English speakers, that rate is below 5%.

For someone writing in Turkish, the situation is even murkier. The agglutinative structure of the language, the possibility of inverted sentence order and the scarcity of formulaic academic phrases all mislead these tools.

2. The lack of training data

Most detection tools built more than 95% of their training data from English texts. Turkish examples are either scarce or absent altogether. The result: the score given for a Turkish text is closer to a random number than to a model’s prediction.

3. The cost of false positives

Even a 1-3% false positive rate for English is problematic for the teacher-pupil relationship. For Turkish, that rate climbs to the 30-60% band. In other words, if you have 30 pupils in your class, the texts of 10 to 18 of them could be wrongly flagged as “written by AI”.

A Turkish example

The paragraph below was written by hand by a Year 10 pupil, without any AI:

“Sanayi İnkılabı Avrupa’da büyük değişimlere neden olmuştur. Üretim biçimi değişmiş, fabrikalar kurulmuş ve şehirleşme hızlanmıştır. Bu süreç sadece ekonomik değil sosyal yapıyı da derinden etkilemiştir. İşçi sınıfının ortaya çıkması yeni toplumsal sorunları beraberinde getirmiştir.”

When pasted into ZeroGPT, the result: 73% AI generated.

We then had ChatGPT write the same paragraph. ZeroGPT’s result: 89% AI generated.

The gap between the two is 16 percentage points. That check is nowhere near enough to draw a decision boundary between the genuine and the artificial. What is more, 73% is already above the “violation” threshold.

A practical approach for teachers

The score a detection tool gives is not evidence on its own. A teacher needs other indicators to hand.

Follow the writing process. Google Docs’ “Version history” feature or Word’s “Track Changes” mode shows whether the pupil actually wrote the text. There is a marked difference between a text pasted in one go and a text built up over time.

Have the topic discussed in class. If a pupil can defend the content of their text orally, the text is theirs. If they cannot, the source is questionable.

Regulate AI use rather than banning it. A policy that can say “you may use ChatGPT at this stage, for this purpose, but the final text must be your own” is healthier than covert use. This is also the core approach of the YAZEK guide: not banning AI, but defining responsible use.

Never make a detection tool’s score the sole grounds for discipline. Accusing a pupil on the basis of a 73% AI score damages the trust built in the classroom beyond repair.

Sources:

  • Liang, W. et al. (2023). GPT detectors are biased against non-native English writers. Patterns, Cell Press.

  • Official documentation of ZeroGPT, GPTZero, Turnitin and Originality.ai (2025).

  • Ministry of National Education (MoNE) (2026). Ethical Guide to Artificial Intelligence Applications in Education.