I(5) The Future of Academic Writing in the Age of Artificial Intelligence

Academic writing is entering a new phase of technological transformation, but the relationship between scholarship and technology is not new.

For decades, researchers have used computers, digital archives, reference management systems, and various software tools to support the creation and evaluation of academic texts. Artificial intelligence (AI) represents a continuation of this long process. The difference is that previous tools mainly assisted human actions. The widespread use of large language models  (LLMs) has already changed how researchers write papers, prepare reviews, summarize literature, and communicate scientific ideas. The question is no longer whether artificial intelligence will become part of academic work. It already has.

The more important question is: what will academic writing become when both humans and machines participate in the production of knowledge?

Many researchers now use AI systems as writing assistants. They ask them to improve sentences, restructure arguments, suggest references, review manuscripts, or evaluate the clarity of academic texts. In this sense, AI has become a new layer in the scholarly process. It is not simply a tool for correcting grammar anymore. It is more than that, for sure. It increasingly participates in the organization and development of ideas.

This creates a new and somewhat unusual situation.

Academic texts produced with the assistance of AI may later become part of the broader information environment used to develop future generations of language models. Conversations between researchers and AI systems, when collected and used under appropriate conditions, may influence how future models understand academic language, argumentation, and scientific communication. The boundary between human generated and machine assisted knowledge production may therefore become increasingly complex.

However, the existence of AI generated text does not mean the disappearance of human authorship.

A text is not only a sequence of well connected sentences. Academic writing is also a decision about what matters. It requires curiosity, judgment, theoretical perspective, and responsibility. These elements remain human. For now.

As Bender et al. (2021) emphasize, language models generate text by identifying patterns in large datasets rather than by possessing understanding or independent intentions. Their ability to produce convincing academic language should not be confused with scientific agency.

This distinction is essential.

An LLM can help a researcher write an article. It can suggest possible arguments. It can even identify weaknesses in a manuscript. But it will never independently enter the academic world as a researcher. It will never send a message saying: “I have written a paper about this topic. Would you like to become the author?”

The initiative remains human.

The future of academic writing will probably not be a competition between humans and machines. It will be a collaboration in which the roles are different. AI systems may become extremely advanced assistants, capable of processing enormous amounts of information and supporting complex intellectual tasks. Humans, however, will remain responsible for asking meaningful questions, interpreting results, and deciding whether an idea contributes something valuable to knowledge.

This transformation will also affect peer review. AI tools may assist reviewers by identifying methodological problems, language issues, missing information, or inconsistencies. Nevertheless, academic evaluation is not only a technical process. A good review requires understanding the importance of an idea, the originality of an approach, and the contribution of a work to a specific field.

Van Dis et al. (2023) argue that AI systems such as ChatGPT create both opportunities and risks for research. Their usefulness depends largely on responsible human oversight and clear academic standards.

Therefore, future universities and research institutions will need new principles for transparency. Researchers may need to disclose how AI tools were used, especially when they contribute significantly to writing, analysis, or evaluation. The central issue will not be whether AI was used, but how it was used.

The future scholar may write differently. The future paper may be produced through a dialogue between human and machine. The process of writing may become less individual and more interactive.

But the final intellectual responsibility will remain with the human author.

Ideas need owners. Arguments need defenders. Knowledge needs people who believe that a question is worth asking.

Artificial intelligence may become a powerful partner in academic writing. It may accelerate discovery and reshape scholarly communication. Yet the direction of knowledge will continue to depend on human imagination, human judgment, and, especially, human curiosity.

Or we need weekness in the text, so we convince everybody that the text is entirely human-written?


References

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. https://doi.org/10.1145/3442188.3445922

Van Dis, E. A. M., Bollen, J., Zuidema, W., van Rooij, R., & Bockting, C. L. (2023). ChatGPT: Five priorities for research. Nature, 614, 224–226. https://doi.org/10.1038/d41586-023-00288-7


Ahmed, O. (2026). The Future of Academic Writing in the Age of Artificial Intelligence. Notes on Language and Artificial Intelligence, I(5). http://turkoloji.net/notes/i5-the-future-of-academic-writing-in-the-age-of-artificial-intelligence

I(4) ÜSKÜP TÜRK AĞZINDA “X GİBİ” YAPILAR VE STANDART TÜRKİYE TÜRKÇESİNDEN ANLAM SAPMALARI

Üsküp Türk ağzında “X gibi” yapısı, standart Türkiye Türkçesindeki benzetme işlevinin ötesine geçerek kalıplaşmış anlam genişlemeleri ve yerel semantik yoğunlaşmalar üretir. Bu yapı, yalnızca benzetme kurmakla kalmaz, aynı zamanda topluluk içi değerlendirme, niteleme ve kültürel kodlama işlevi de üstlenir.

Bu ağızda aşağıdaki örnekler belirli anlam kalıplaşmalarını göstermektedir.

[1] “güna gibi”
Standart Türkçede olsaydı, “günah gibi” olurdu. Bu ifade anlamsal kayma ile “acınacak durumda, zayıf ve kötü halde” anlamına sabitlenmiştir.
“Oturmiş orda güna gibi.”

[2] “rak gibi”
Makedonca “рак”/”rak” (yengeç) etkisiyle gelişmiş olup renk yoğunluğu ve aşırılık bildiren bir pekiştirme işlevi kazanarak “aşırı derecede kırmızı” anlamında kullanılır.
“Yüzi güneşten rak gibi olmiş idi.”

[3] “Nasradin gibi”
Nasrettin Hoca figürünün yerel yeniden yorumlanmasına dayanır ve “kurnaz, uyanık ve pratik zekâsı güçlü kişi” anlamını taşır.
“Kenardan bakaydi Nasradin gibi.”

[4] “çarık gibi”
Standart dilde somut nesne benzetmesi iken yerel kullanımda “esnek, kolay şekil alan fakat kopmayan yapı” anlamına dönüşmüştür.
“Bu ekmek çarık gibi bayat.”

[5] “yag gibi”
“Tereyağı gibi” anlamının genişlemiş biçimi olup özellikle erkeklerde takım elbise bağlamında “düzgün, pürüzsüz ve estetik biçimde duran (bir takım elbise)” anlamında kullanılır.
“Üstündeki rubalar yag gibi.”

[6] “ilân gibi”
Standart “yılan gibi” benzetmesinin yerel fonetik uyarlamasıdır ve “hızlı, çevik ve ele avuca sığmayan” anlamını ifade eder.
“Kütek yeecek, kaçaydi ilân gibi.”

[7] “Draçova gâuri gibi”
“Draçova gâvuru gibi” yerel sosyo-kültürel algıya dayalı stereotipik bir yapı olup “çok kötü, kaba, pis, uygunsuz ve dinsiz” anlamına sabitlenmiştir. “Draçova”, Üsküp’te bir semtin (eskiden bir köyün) adıdır.
“Kirli Draçova gâuri gibi.”

Genel olarak Üsküp Türk ağzındaki bu yapılar, “gibi” benzetme edatının semantik sınırlarını genişleterek, yerel kültürel deneyim, çok dilli temas ve tarihsel hafızanın birleşimiyle anlamı kalıplaştıran özel bir ifade sistemi oluşturur. Bu durum, Balkan Türk ağızlarında görülen yoğun temas kaynaklı anlamların yeniden yapılandırmanın tipik bir örneğini teşkil eder.


How to cite this post:

Ahmed, O. (2026). Üsküp Türk Ağzında “X gibi” Yapılar ve Standart Türkiye Türkçesinden Anlam Sapmaları. Notes on Language and Artificial Intelligence, I(4). https://turkoloji.net/notes/i4-uskup-turk-agzinda-x-gibi-yapilar-ve-standart-turkiye-turkcesinden-anlam-sapmalari

I(3) AI Systems and the Limits of Bilingual Sentence Interpretation: A Linguistic Perspective

Artificial intelligence systems that process natural language have now achieved remarkable performance in monolingual tasks. However, this is not the case under all circumstances. Their capacity to interpret bilingual or mixed language sentences remains structurally constrained. This is especially true when it comes to the Macedonian Turkish dialects that fall within the boundaries of the Balkan Sprachbund. This limitation is not primarily a matter of vocabulary size or training data volume. It concerns how meaning is represented in statistical language models. Small and underrepresented languages are even more affected.

A useful set of examples, collected by the author in the Turkish dialect of Skopje, can be observed in mixed Turkish and Macedonian influenced varieties, particularly in related urban bilingual environments:

[1] “Bügün gittım snimanyeye Televiziyada.”

In its intended interpretation, this sentence refers to attending a television recording session today. The lexical item “snimanye” corresponds to the Macedonian noun “снимање,” meaning “recording,” while “televiziyada” refers to a television broadcasting institution or studio context, with locative case marking at the end. Locative case marking is often used as dative case in the Turkish dialects in North Macedonia (Ахмед, 2004, p. 57). A human bilingual speaker reconstructs the meaning by integrating Turkish morphosyntax with Macedonian lexical insertions and pragmatic knowledge of media production environments.

Additional examples from Skopje Turkish dialect usage further illustrate this phenomenon:

[2] “Klikala te buni.”

Intended meaning: “Click exactly this one.” This reflects a hybrid structure where a Macedonian derived verb form is integrated into a Turkish communicative imperative context.

[3] “Maksıma ver tsırtalasın, em bitti dovasi.”

Intended meaning: “Give a child to draw, and that’s it.” Here, “tsırtalamak” derives from Macedonian “црта” (to draw) verb, integrated into a Turkish verbal framework, while the clause structure reflects conversational compression typical of bilingual speech.

Verbs “klikalamak” and “tsırtalamak” given in the examples [2] and [3] are copied verbs (Ahmed, 2016). Sentences of this type are not consistently and accurately translated by currently available large language models.

These examples demonstrate that bilingual speech in contact zones is not random lexical mixing but a structured communicative system shaped by long term language contact, cognitive economy, pragmatic inference and sense of being part of a society.

Current AI systems typically process such input through token segmentation and probabilistic association. While they may approximate meaning under favorable conditions, they do not reliably reconstruct the intended cross linguistic mapping when orthography is inconsistent or when lexical boundaries shift across languages within a single clause. As Grosjean (1982) argues in his foundational work on bilingualism, bilingual speech is not a mixture of two monolingual systems but a fully integrated communicative mode. This integration poses a structural challenge for models trained primarily on monolingual distributions or artificially segmented multilingual corpora.

From a computational perspective, modern large language models are built on the transformer architecture, which processes token sequences through learned statistical regularities via self attention mechanisms (Vaswani et al., 2017). Although multilingual pretraining improves robustness, it does not guarantee stable interpretation of hybrid forms where phonetic spelling, code switching, and pragmatic compression occur simultaneously. Zhang et al. (2023) found that multilingual large language models are not yet competent code switchers.

Muysken (2000) similarly demonstrates that code switching is rule governed and context sensitive, not random mixture.

This implies that correct interpretation requires not only lexical alignment but also discourse level inference, which remains an area of partial competence for current AI systems. Newer evaluation work supports this directly. Mohamed et al. (2025) note that existing benchmarks concentrate on surface level tasks such as language identification, sentiment, and part of speech tagging, leaving deeper semantic and reasoning capacities largely untested. Sheth et al. (2025) catalog ongoing evaluation efforts across more than 300 studies and identify open problems that remain unresolved despite continued progress in multilingual pretraining. These findings, drawn from evaluations published in 2025, reinforce the original argument empirically. The structural limitation is not primarily about scale or vocabulary coverage. It concerns how statistical models represent meaning when two linguistic systems are genuinely fused rather than alternated, a distinction that matters directly for dialect contact zones such as the Balkan Sprachbund.

In all the provided examples, a human bilingual speaker reconstructs meaning through a combination of phonological approximation and shared cultural context. And pragmatic inference, too. The system identifies “snimanye” with recording contexts, “tsırtalamak” with drawing activity, and interprets mixed imperatives such as “Klikala te buni” within a situational frame of digital interaction.

The limitation observed here is not the absence of bilingual data in training, but the absence of grounded pragmatic interpretation that dynamically integrates cross linguistic cues in real time. As a result, AI systems may produce plausible monolingual interpretations while missing the intended bilingual communicative act.

This distinction is critical for applications in low resource language contexts, dialectal variation, and informal digital communication, where hybrid linguistic forms are common. It also suggests that future improvements in AI language understanding will require deeper integration of discourse modeling, pragmatic inference, and structured representations of bilingual speech behavior.


References

Ахмед, О. (2004). Морфосинтакса на турските говори од Охридско Преспанскиот регион. Необјавена докторска дисертација. Филолошки факултет „Блаже Конески“, Универзитет „Св. Кирил и Методиј“, Скопје.

Ahmed, O. (2016). Copied verbs in Turkish dialects of Macedonia. In E. Á. Csató, B. Karakoç, & A. Menz (Eds.), The Uppsala Meeting: Proceedings of the 16th International Conference on Turkish Linguistics (pp. 9–18). Harrassowitz Verlag.

Mohamed, A., Zhang, Y., Vazirgiannis, M., & Shang, G. (2025). Lost in the mix: Evaluating LLM understanding of code-switched text. arXiv. https://arxiv.org/abs/2506.14012

Sheth, R., Sinha, S. R., Patil, M., Beniwal, H., & Singh, M. (2025). Beyond monolingual assumptions: A survey of code-switched NLP in the era of large language models across modalities. arXiv. https://arxiv.org/abs/2510.07037

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30 (pp. 5998–6008).

Zhang, R., Cahyawijaya, S., Cruz, J. C. B., Winata, G., & Aji, A. F. (2023). Multilingual large language models are not (yet) code-switchers. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 12567–12582). Association for Computational Linguistics.


How to cite this post:

Ahmed, O. (2026). AI systems and the limits of bilingual sentence interpretation: A linguistic perspective. Notes on Language and Artificial Intelligence, I(3). https://turkoloji.net/notes/i3-ai-systems-and-the-limits-of-bilingual-sentence-interpretation-a-linguistic-perspective/

 

 

I(2) Why AI-Text Detectors Do Not Work and Why Their Results Cannot Be Trusted

With the expansion of artificial intelligence (AI) and the increasing use of AI-generated text, teachers, publishers, and other professionals have begun to face the problem of determining whether a text is AI-generated, human-written, or the result of a combination of both.

AI-text detectors are tools that claim to distinguish between human-written and machine-generated text. Or are they really what they claim to be?

Top academic institutions, employers, and publishers have begun adopting them as gatekeeping instruments. The research evidence, however, shows that these tools are unreliable in ways that make their use actively harmful.

The core problem behind AI-text being flagged is technical.

Strip away the marketing language and what these detectors actually measure is adherence to standard written English. Text that follows grammatical rules and formal organizational conventions scores high for AI probability. Text with colloquialisms, fragments, and informal phrasing scores low.

The tools are not detecting AI. They are detecting writing quality, and penalizing students for demonstrating exactly the skills they were taught. Meaning that if you don’t want to be flagged, you should use worse language with a lot of colloquialisms and slang. Wow!

This flaw is demonstrable by testing the tools against material that predates the existence of generative AI entirely. As I discuss in my book (Ahmed, 2026, pp. 69-72), three short stories written in 1995 on a word processor, with no AI assistance because no such assistance existed for creative writing at the time, were fed through a leading detector. All three were flagged at over 80 percent AI probability. A system that cannot distinguish between human writing from 1995 and AI writing from 2026 is not measuring what it claims to measure.

I use AI to check my English grammar. I mostly accept AI’s suggestions for better sentence structures, but then that text is flagged as AI-generated. Sometimes hybrid. Sometimes even 100%.

Why? Because I used AI for grammar check. Didn’t we have that for decades in other forms, such as Grammarly? What about editors? Aren’t those people correcting your text if needed?

Other researchers have replicated this pattern. Classic essays, published novels from the 1960s, and historical documents written long before computers existed are routinely flagged by commercial detectors.

The reason is consistent: well-written, formally structured text produces the same statistical patterns whether it was written by a person thirty years ago or generated by a model today, because the models were trained on human writing and learned its patterns.

The false positive problem is particularly severe for non-native speakers of English. Like me. Liang et al. (2023) evaluated seven widely used GPT-detection systems using two datasets: essays written by native English-speaking US eighth-grade students and essays written by non-native speakers for the TOEFL exam. The detectors correctly classified most of the native-speaker essays, with a mean false positive rate of 5.1 percent. For the TOEFL essays, the mean false positive rate was 61.3 percent, and all seven detectors unanimously flagged 19.8 percent of those human-written essays as AI-generated. The study concluded that non-native speakers’ writing tends toward lower lexical richness and syntactic diversity, which the detectors read as machine output. These tools embed and amplify a structural bias against multilingual writers.

OpenAI itself released an AI classifier in January 2023 and withdrew it on July 20, 2023, citing a low rate of accuracy. The company’s own documentation stated that the classifier correctly identified AI-written text only 26 percent of the time while incorrectly labeling human-written text as AI-generated 9 percent of the time (OpenAI, 2023). A 9 percent false positive rate applied to a classroom of 30 students means roughly three students face wrongful accusation on any given assignment.

The perverse incentive that follows from this is real and documented. Students learn quickly that formal, well-structured writing is more likely to be flagged. The rational response is to degrade the work: include sentence fragments, add casual phrasing, introduce minor errors. Institutions are inadvertently teaching students that excellence is dangerous and that the goal is plausible mediocrity rather than genuine quality. Is this progress or regression? That’s what we have to think about.

The business model behind the detection industry compounds the problem. Several companies sell both detection services and “humanization” services that rewrite AI-generated text to pass detection. That is the main point. To sell you fear and to take your money. A company that profits from flagging text has a financial incentive to flag aggressively, producing more worried users and more demand for the paid solution. The humanization process itself uses AI to fool the AI detector, which demonstrates conclusively that the underlying signal the detectors rely on is not stable or meaningful.

There is also a fundamental epistemological problem. Running the same text through multiple detection tools produces wildly divergent results. One tool may return 85 percent AI probability while another returns 23 percent on the same passage. These are not measurement errors around a true value. There is no ground truth to converge on, because the tools are applying different statistical models to a problem that has no clean solution.

Better approaches to academic integrity exist and do not require these tools. Teachers who know their students can identify suspicious work through patterns no algorithm captures: a sudden unexplained jump in sophistication, inconsistency with the student’s voice in class discussion, or a polished final product with no visible process of development. Oral assessment, where students present and defend their work, exposes misuse immediately. Process documentation, including drafts and outlines, provides evidence of genuine engagement that cannot be faked by submitting a single AI-generated output.

Accepting detector output as evidence in disciplinary proceedings means accepting a tool with a documented and significant error rate, a demonstrated bias against non-native speakers, a conflicted commercial incentive to over-flag, and no methodology that has been independently validated against real-world academic writing.

In legal and scientific reasoning, a measurement instrument must demonstrate validity before its outputs are used to make consequential decisions about people. AI-text detectors have not met this standard. Their continued use in contexts that carry real stakes for real individuals is not justified by the available evidence.


References

Ahmed, O. (2026). If humans can learn from books, why can’t AI? Reconsidering the training data debate. Intelligentia Nova Press.

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779

OpenAI. (2023). New AI classifier for indicating AI-written text. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/


How to cite this post:

Ahmed, O. (2026). Why AI-text detectors do not work and why their results cannot be trusted. Notes on Language and Artificial Intelligence, I(2). https://turkoloji.net/notes/i2-why-ai-text-detectors-do-not-work-and-why-their-results-cannot-be-trusted/

I(1) “YORUM YORUM KONUŞMAK” VE “İSTİLÂİCE”

Standart Türkiye Türkçesinde şimdiki zaman eki /-yor/ biçimindedir. Bu ek, geniş bir kullanım alanına sahip olup konuşma dilinde de temel görünümü temsil eder.

Üsküp Türk ağzında ise bu ek standart biçimiyle kullanılmaz. Bunun yerine farklı varyantlar ve kiplik yapıları tercih edilmektedir (bkz. Ahmed, 2014).

Türkiye Türkçesi konuşurlarının Makedonya Türkleriyle iletişiminde ve standart Türkçenin öğrenim süreçlerinde bu farklılık belirgin bir algısal ayrışmaya yol açmaktadır. Bu bağlamda “yorum yorum konuşmak” şeklinde yerel bir ifade, standart /-yor/ yapısının yoğun ve yabancı algılanan kullanımına yönelik bir gözlemi yansıtmaktadır.

Bu durum, şimdiki zaman kategorisinin yalnızca morfolojik bir karşıtlık değil, aynı zamanda sosyodilbilimsel bir varyasyon alanı olduğunu göstermektedir.

Örneğin:

[1] O da çok yorum yorum konuşi. [O da abartılı bir şekilde standart Türkçeyi konuşuyor.]

[2] Bir afta İstanbol’da em ben da yorum yorum başladım konuşam. [Bir haftadır İstanbul’dayım ve ben de artık İstanbullular gibi standart Türkçeyi konuşmaya başladım.]

[3] Kırdi dilıni yorum yorum konuşmaktan. [Standart Türkçeyi konuşmaktan dilini kırdı.]

Üsküp Türk ağzında “standart Türkçe” ifadesi, İstanbul ağzı anlamında kullanılmaktadır.

Bunun yanında, standart Türkçe için “istilâî” veya “istilâice” gibi adlandırmaların da kullanıldığı görülmektedir. 2000’li yıllardan sonra bu kullanımların neredeyse tamamen ortadan kalktığı gözlemlenmektedir. Günümüzde bu ifadeye kısmen Arnavut kökenli konuşurlar arasında rastlanabilmektedir.

[4] Dili sanki istilai idi. [Kullandığı dil İstanbul ağzını andırıyordu.]

[5] Geldi birisi, istilaice konuşidi. [Biri geldi ve İstanbul ağzını (standart Türkçeyi) konuşuyordu.]

Büyük bir ihtimalle, bu ifade ilk olarak Osmanlı döneminde Arnavut kökenli unsurlar tarafından kullanılmış olup, zamanla yaygınlaşmıştır.

Oktay Ahmed


Kaynakça

Ahmed, O. (2014). Üsküp Türk Ağzında Kip Ekleri. TÜRÜK Uluslararası Dil Edebiyat ve Halk Bilimi Araştırmaları Dergisi, 1(4), 1–36. https://dergipark.org.tr/en/pub/turuk/article/162156


How to cite this post:

Ahmed, O. (2026). “Yorum Yorum Konuşmak” ve “İstilâice”. Notes on Language and Artificial Intelligence, I(1). https://turkoloji.net/notes/i1-yorum-yorum-konusmak-ve-istilaice/