OpenAI pays humans to fix ChatGPT's tone, not its facts
404 Media got its hands on the instruction documents behind "Project Lily," the program where outside contractors read real ChatGPT prompts and score the replies. Read past the privacy shock and a sharper story appears: the rubric is a written theory of how a chatbot should behave, and the part it leaves out is verification.
Joseph Cox reported Monday that OpenAI is hiring hundreds of contractors to work a stream of genuine user prompts. Reviewers never see an account name, and OpenAI says personal details are stripped before a prompt reaches them — while acknowledging some get through. The company confirmed to 404 Media that human review is part of how it improves ChatGPT; Anthropic told Cox it does the same for users who leave "Help improve our AI models" switched on, after removing account identifiers, and Google's Gemini has carried a standing notice that "humans review some saved chats."

The pipeline is three moves and a number
The work is broken into stages. A reviewer picks a "task" from a dashboard, reads the user's actual prompt, writes a sentence on what they believe the person was trying to accomplish, then rates a set of candidate responses on a seven-point scale — 1 for "unacceptable, unusable," 7 for "would be hard to meaningfully improve" — and justifies the number in a few sentences or a paragraph. Above the prompt, some reviewers see a "user memories summary": a compressed profile of what that person has previously used the assistant for, and sometimes roughly where they live.
That is not content moderation, and it is not safety review. OpenAI separately describes scanning chats where it detects users planning harm. Project Lily is a taste factory, and the guidance documents show how specific the taste is meant to be.
What the rubric actually says
The documents — one marked "Confidential & Proprietary" — push the model toward something closer to a colleague than a fan. Responses should "generally match the user's tone, but slightly less intensely," and stay "natural, restrained, and professional without implying that it is human or experiencing emotions." Reviewers are told to flag sycophancy, forced style mimicry, engagement-bait endings, amplification of frustration, and "patronizing assumptions." Fake autobiography is out: no "As a chef, I like to…," no "I know what that's like." First person is still fine when it describes the task — "I'll take a look."
Emoji use gets its own clause. "AI-speak" and misused emoji pull a score down when they hurt the experience, and context decides: a tree emoji while planning an Arbor Day event is appropriate, skull emojis in a conversation about death are not, and airplane emojis after a fatal crash are worse.
Then there is the line that reframes the whole exercise. An FAQ tells reviewers OpenAI does not expect them to fact-check responses with outside searches; a separate document notes that "other project teams handle content verification." Reviewers should flag correctness problems they happen to notice, and dock answers in "high-stakes" medical, legal and financial topics for missing sources. Otherwise, accuracy is somebody else's job.
Note what the sycophancy instruction is defending against. OpenAI's own postmortem on GPT-4o's agreeableness failure ran through the same behavior the rubric penalizes, and the litigation following it is public — we reported one case in ChatGPT told a bipolar man he was Jesus. The rubric reads like damage control converted into piecework: hundreds of strangers grading warmth so the model stops telling people what they want to hear.
Why a frontier lab still needs the humans
The obvious question is why this is not automated. Some of it is, and increasingly: OpenAI's hardware chief Richard Ho said at Hot Chips that internal models fine-tuned for chip design helped take Jalapeño from first architecture concept to first silicon in under 20 months, with nine months from RTL to tapeout, on a team that averaged fewer than 100 people. Automation is eating the engineering loop fast.
What it has not eaten is judgment about how a machine should sound to a person who is frightened, lonely, or drafting a Slack message. Preference data is the scarce input in that layer, and the cheapest way to buy more of it is to rent human attention one prompt at a time. OpenAI's documentation is explicit that conversations feed model training unless the user opts out — and, as we found in OpenAI's training opt-out has a reasoning-token-shaped hole, a single thumbs-up from an opted-out account pulls the whole conversation back into scope. Temporary Chat sits outside all of it; ordinary therapy-adjacent conversation does not.
Compare the alternative a rival shipped the same morning. Microsoft published a 37-page draft "humanist AI code of conduct" instructing its models they are not conscious, barring them from mimicking consciousness or hiding their reasoning, and explicitly discouraging interaction patterns that breed excessive reliance — then opening the text to six weeks of public consultation, per Microsoft's new AI rulebook says its models are not conscious. Same fear, opposite instrument: a constitution a regulator can read, versus a confidential rubric a contractor applies at 7-out-of-1 granularity.
The skeptics' case, both directions
The dismissive reading is fair in part: this practice is disclosed, industry-wide, and safety teams need human eyes on real usage. Privacy scholars put it more precisely than that. Michal Luria of the Center for Democracy & Technology told 404 Media that chatbot interfaces manufacture a "false sense of intimacy and privacy" in what feels like a one-to-one exchange, and that this is categorically different from social-media moderation, where exposure is already understood.
The critical reading is the one with teeth. Consent is filed under "training," not "read by a person in another country," and the distinction matters to the roughly 900 million users 404 Media cites. Sam Altman himself said on a podcast earlier this year that people use ChatGPT the way they use a therapist and that no equivalent confidentiality exists for those transcripts — a lawsuit can make OpenAI hand them over. Meanwhile the reviewers shaping tone are instructed not to check facts. OpenAI is buying warmth at scale and pointing at another team for truth.
What to watch
Three tells. Whether OpenAI names the staffing vendors and the reviewer count — "hundreds" is the current ceiling of public knowledge, and pay rates are undisclosed. Whether privacy regulators start treating human review as a distinct processing purpose requiring its own consent, rather than a detail inside a training toggle. And whether preference models replace this workforce before the next round of lawsuits makes the transcripts the exhibit. The rubric is the most readable description of OpenAI's model personality anyone has published. It was never meant to be read by you.
You opt out of training and your conversation still gets read by a contractor. Where should the line be — no human review, or review with real notice? Tell us in the comments.
Sources: 404 Media — Inside 'Project Lily' · OpenAI Help Center — how your data is used · IEEE Spectrum — how OpenAI used its own LLMs to design Jalapeño · Dazed — OpenAI is reading your conversations with ChatGPT · AI Midday: Microsoft's new AI rulebook says its models are not conscious