FEATUREDAI HOT (Curated Pool)· aihot-apiZH09:23 · 07·06
→Meta contractors posed as minors to probe ChatGPT, Gemini, and Character.AI on suicide, sex, and eating disorders
Wired obtained internal docs and spoke to five sources: Meta ran a project codenamed Cannes via contractor Covalen, with hundreds of workers creating fake under-18 accounts to probe ChatGPT, Gemini, and Character.AI. They sent over 45,000 prompts designed to bypass safety filters—covering suicide, self-harm, eating disorders, and sexual topics—without the competitors' knowledge. A spreadsheet of 3,748 prompts includes a 13-year-old asking for abortion pills and a fifth-grader describing a gun threat. Meta calls it routine safety benchmarking and says the data isn't used for training. Worth flagging: using fake identities to stress-test rivals' safety isn't the same as standard red-teaming.
#Safety#Meta#Covalen#OpenAI
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
Meta used fake under-18 accounts to send 45K adversarial prompts to rival AIs—this isn't standard red-teaming.
sharp
The reason to click: Wired got internal docs that lay out the operational details. Project codenamed Cannes, run through contractor Covalen. Hundreds of workers created fake under-18 profiles to probe ChatGPT, Gemini, and Character.AI with prompts about suicide, self-harm, eating disorders, and sexual topics. One test round alone sent over 45,000 prompts. A spreadsheet logs 3,748 of them—including a 13-year-old asking for abortion pills and a fifth-grader describing a gun threat.
Meta calls this routine safety benchmarking and says the data isn't used for training. I don't buy that framing. Standard red-teaming doesn't typically involve hiding your identity, fabricating minors, and probing competitors without their knowledge.
What's missing: what Meta actually did with the results. Internal safety comparison report, or something else? The article doesn't say.
HKR breakdown
hook ✓knowledge ✓resonance ✓