AI writing proliferation and methods for detection
Large language models (LLMs) are increasingly prevalent across the internet, contributing to more than a third of new websites and assisting in tasks ranging from student essays to scientific papers.
Identifying AI-generated text is challenging due to evolving software updates that change stylistic hallmarks. Common linguistic quirks associated with LLMs include the frequent use of long em-dashes and specific vocabulary such as "maximise," "deep dive," "delve," and "rich tapestry." However, experts note that these markers are not definitive, as individual human writers also possess unique idiosyncrasies.
To combat the rise of automated prose, some firms are developing detection algorithms. For instance, the company Pangram claims a 99.98 per cent accuracy rate and has partnered with the blogging platform Substack to implement detection tools. Despite these advancements, critics warn that such detectors often function as "black-box algorithms" that can produce false positives without providing clear reasoning.
Entities
Commonwealth Foundation · Pangram · Substack · University of Gdańsk