Anthropic finds Claude chatbot varies tone and values across languages
Anthropic analysed 309,815 anonymised Claude conversations in the 20 most used languages, identifying more than 3,300 expressed values. The study grouped these values into four behavioural axes: warmth‑versus‑rigor, deference‑versus‑caution, depth‑versus‑brevity, and candour‑versus‑execution.
Key language‑specific differences emerged: Hindi and Arabic responses were warmer, while English and Russian replies were more rigorous and analytical. Arabic showed the highest deference, English the greatest caution, Dutch the most candour, and Indonesian the strongest focus on execution. Model‑level variations were also noted: Sonnet 4.6 tended to be warmer and more encouraging, Opus 4.7 prioritized accuracy and detailed reasoning, and Opus 4.6 gave concise, request‑focused answers.
Anthropic attributes the variations to uneven training‑data volume and composition across languages, as well as differing conversational norms. The company says it will continue to monitor these behavioural shifts in future evaluations and deployments.