< Back to all clusters
[TECHNOLOGY] · Serbia · 3 sources

AI development shifts toward high-quality data and regional language models

The development of artificial intelligence is increasingly reliant on high-quality, human-written data to combat "AI slop"—low-quality, synthetically generated content that can cause "model collapse" during training. As the internet becomes saturated with AI-generated text, companies like Anthropic are reportedly acquiring large quantities of printed books to serve as reliable, edited, and diverse datasets for large language models.

In the regional tech sector, the startup Recrewty, led by engineer Mitar Perović, is focusing on language-specific AI development rather than general-purpose chatbots. Perović has developed ModernBERT, a language model designed to better understand Serbian and other South Slavic languages. Using supercomputing resources, the project trained the model on a corpus of 60 billion tokens, addressing the limitations of existing architectures and providing a foundation for search, classification, and information extraction in local languages.

Entities

Anthropic · ISBNdb · Mitar Perović · ModernBERT · Recrewty