< Back to all clusters
[TECHNOLOGY] · 3 sources

started · updated

AI technology research reveals security flaws and detection biases

Researchers have developed a method to detect security vulnerabilities in AI-generated code by analyzing the internal activations of open-weight large language models. Using linear probes, the team achieved a 61-67% success rate in identifying vulnerable functions, outperforming the models' own prompted assessments. This approach reveals security signals within the model's internal workings that are not apparent through direct querying.

Separately, studies on AI writing detectors have revealed a significant fairness gap. Testing on TOEFL essays showed that detectors wrongly flagged an average of 61% of genuine writing by non-native English speakers as AI-generated. These tools often mistake the predictable phrasing used by multilingual students for machine-made text. Due to these high false-positive rates, institutions like Vanderbilt University have moved away from using automated detection tools for student assessments.

Entities

MAAS EdTech · Turnitin · Vanderbilt University