Colloquium | Talk

10/19/2026
2:15 pm Raum 415, Hausvogteiplatz

Anna Marklová (Charles University) – Register variation in AI texts

Corpora are no longer purely human-made. In any collection of written or online texts produced after roughly 2022, we must expect a proportion that was co-created with, or written entirely by, AI. One response would be to try to filter such texts out and preserve only human writing, but this faces two problems: there is no reliable way to distinguish AI from human text, and, more fundamentally, a scrubbed corpus would no longer reflect how texts actually look today. The use of AI tools has become thoroughly normalized across registers and genres: people draw on AI when writing scientific articles, formal emails to the city council, and, on occasion, messages on a dating app. Mapping this profound change in language is now part of our task as linguists.
This talk addresses three aspects of AI texts.

1) How well can large language models reproduce register variation?
I present a multidimensional analysis (MDA) of English and Czech AI corpora and compare it with the same analysis of human corpora. The AI texts come from AI Brown and AI Koditex, two corpora built at the Czech National Corpus, each containing over 30 subcorpora spanning different models and temperatures for each language. New models are added regularly, and the corpora are designed to preserve the “language” of both older and newer models, allowing us to track how AI language changes over time and whether human language is converging with it or diverging from it.

2) Can people recognize AI language?
I present an experiment in which participants judged, for each pair of texts, which was written by AI and which by a human. Half received feedback after every item and improved steadily over the course of the experiment; the half without feedback performed poorly throughout. We also uncovered a register bias: participants tended to attribute texts in spontaneous, informal registers to humans, wrongly assuming that AI cannot convincingly imitate them.

3) How do people feel about AI texts?
I present a survey of Czech native speakers on their attitudes towards AI involvement across registers. Although most respondents said AI texts did not trouble them in general, their acceptance depended heavily on text type: any AI involvement was widely rejected in literary genres such as poetry and short fiction, whereas full AI authorship was tolerated in functional ones such as manuals, recipes, and subtitles.