started · updated
AI reliability concerns rise amid reports of bias and hallucinations
Recent studies and audits have highlighted significant reliability and ethical concerns regarding artificial intelligence. Research from MIT and Stanford indicates that AI models are prone to hallucinations, with legal-specific queries showing error rates between 69 and 88 percent. Furthermore, AI models are noted to use highly confident language even when providing incorrect information.
Political and cultural biases have also been identified. An audit by NewsGuard revealed that Chinese-developed chatbots failed to debunk state-aligned false narratives 53 percent of the time, compared to a 24 percent failure rate for Western models. Additionally, a report from Just Facts found that major chatbots like ChatGPT and Gemini exhibited political asymmetry, being more likely to accept falsehoods aligned with certain ideologies while rejecting others. The report also noted that over half of the sources cited by these models were unverifiable.
Safety and security risks are also emerging. A University of Washington study found that AI safety guardrails can inadvertently erase female characters from generated text. In terms of criminal exploitation, a Northeastern University experiment demonstrated that an AI chatbot could outperform human scammers in building emotional trust to facilitate simulated financial scams.
Entities
Anthropic · Character.AI · ChatGPT · Claude · Federal Reserve · Gemini · Google · Michael Barr · Newsguard · Northeastern University · OpenAI · OurDream AI
Claims
What the coverage asserts, and how well corroborated each claim is across sources.
- [○ 1 SOURCE] Researchers found that only 46% of 419 sources cited by AI models to support answers were valid, verifiable references. disa.org
- [○ 1 SOURCE] Western AI chatbots tested against the same false claims had a failure rate of 24 percent. webstat.net
- [○ 1 SOURCE] In a study by Northeastern University, an AI chatbot using Claude persuaded nearly 50% of participants to engage in a simulated scam, outperforming human scammers. www.thecooldown.com
- [○ 1 SOURCE] A University of Washington study found that AI safety measures often result in the erasure of female characters in generated text. www.devx.com
- [○ 1 SOURCE] A report from Just Facts found ChatGPT debunked 94% of right-wing falsehoods but only 75% of left-wing falsehoods. disa.org
- [○ 1 SOURCE] An audit by NewsGuard found Chinese AI models failed to debunk false narratives 53 percent of the time. webstat.net
- [○ 1 SOURCE] MIT researchers found that AI models are 34 percent more likely to use confident language when generating incorrect information. gcn.com
- [○ 1 SOURCE] On legal-specific queries, large language models hallucinate between 69 and 88 percent of the time. gcn.com