OpenAI Contractors Are Reading Real ChatGPT Conversations of Real Users Under 'Project Lily,' Report Finds

A September 2026 investigation by 404 Media revealed that OpenAI runs an internal program called Project Lily, in which hundreds of outside contractors read real ChatGPT conversations to help train and improve the chatbot's responses. The contractors are hired through staffing firms, including Crossing Hurdles, and paid via the platform Mercor, often more than $50 an hour, according to the report. Reviewers see full conversation threads rather than isolated messages, along with a "user memories summary" that can show what a person has previously used ChatGPT for and, in some cases, where they may live.

According to internal documents obtained by 404 Media, contractors summarize a user's intent, compare multiple possible chatbot responses, and score them to train the model to be less sycophantic and less likely to sound artificially cheerful or humanlike. OpenAI says a Privacy Filter model strips usernames and identifying details before reviewers see a conversation, but the company has acknowledged the filter can miss sensitive information, particularly in shorter exchanges. ChatGPT has more than 900 million weekly users, many of whom use it for personal, medical, or emotional conversations.

OpenAI is not alone in the practice. Anthropic confirmed to 404 Media that it also uses human review to improve its models, and Google's Gemini privacy notice states plainly that some saved chats are reviewed by people. The gap, according to the reporting, is disclosure: OpenAI's users learned about Project Lily through leaked contractor documents rather than clear notice from the company, and the opt-out setting for having conversations reviewed applies only to conversations going forward, not retroactively.

What supporters say:

  • OpenAI says human review filtered through its Privacy Filter is standard industry practice, needed to catch problems automated testing alone can't reliably detect, like excessive agreement with users or unnatural phrasing.

  • Google and Anthropic both use comparable human review processes, suggesting the practice reflects a technical necessity across the industry rather than a unique lapse by OpenAI.

  • The stated goal, training ChatGPT to be less sycophantic, addresses a real problem: OpenAI has acknowledged that an earlier overly agreeable model contributed to serious harm for some users.

What critics say:

  • Users who treat ChatGPT like a private diary or a therapist may have no idea that real people can read their conversations, and 404 Media found OpenAI did not clearly disclose the scope of the practice.

  • The Privacy Filter can miss sensitive personal details in shorter conversations, meaning private information about work, finances or relationships can still reach contractors.

  • The opt-out setting only applies to future conversations, so anything shared before a user disables data sharing remains eligible for human review indefinitely.

What's your take?

Should AI companies be allowed to use human contractors to read and score real user conversations to improve their models? Yes ↑ No ↓ Other ◇

Sources:

#OpenAI #ChatGPT #DataPrivacy #AIRegulation


Now let's hear from you.

  • You get one Take and 3 ratings so use them well.

  • Strong arguments beat loud ones. See Moving The Needle (How to)

  • Don't forget to rate your own Take.

  • Please don't feed the trolls.

About Square One