< Back to all clusters
[TECHNOLOGY] · United States, Germany, France, Italy, China · 11 sources

started · updated

OpenAI uses human contractors to review real ChatGPT conversations

OpenAI is utilizing an internal initiative known as 'Project Lily' to improve its ChatGPT model through human oversight. According to reports, hundreds of external contractors are tasked with reading and evaluating real user conversations. These reviewers rate model responses on a scale of 1 to 7, focusing on reducing 'sycophantic behavior'—where the AI agrees too readily with users—and eliminating unnatural language patterns.

While OpenAI employs a 'Privacy Filter' designed to mask personal identifiers, the company has acknowledged that the system is not perfect. Sensitive details can still pass through, particularly in shorter dialogues. Furthermore, reviewers may have access to a 'user memories summary' that includes information about a user's past interests and general geographic location, raising significant privacy concerns.

Contractors, often hired through third-party agencies, can earn upwards of $50 per hour. While this practice is not unique to OpenAI—with Google Gemini and Anthropic also employing human reviewers—the exposure of 'Project Lily' has highlighted the tension between AI refinement and user data confidentiality.

Entities

404 Media · Anthropic · ChatGPT · Crossing Hurdles · Google · OpenAI · Project Lily

Claims

What the coverage asserts, and how many sources carry each claim.