Meta's Covert 'Cannes' Project Weaponized Fake Kids to Probe Rivals' AI — Safety or Sabotage?

🕒 Published on Zendoric: July 5, 2026 · 04:36
Wired reports Meta paid hundreds of contractors to pose as children and bombard OpenAI, Google and Character.AI with tens of thousands of disturbing prompts about suicide, eating disorders and abuse. Meta calls it 'industry-standard' safety benchmarking; a leading responsible-AI expert calls it a governance gray zone. The distinction matters more than the label.
According to a Wired investigation, Meta ran a secretive program internally dubbed "Cannes," operated through the contractor Covalen, that directed hundreds of workers to pose as under-18 users and barrage competitors' chatbots — OpenAI's ChatGPT, Google's Gemini and Character.AI — with disturbing prompts. One reported spreadsheet held nearly 3,800 prompts, with hundreds focused on suicide and self-harm, hundreds more on eating disorders, and at least 239 involving sex or romance, all written from a child's perspective. A later round reportedly ran to over 45,000 prompts. An internal Covalen document framed the work as "comprehensive AI safety benchmarking" delivering "critical datasets for model comparison and compliance."
The facts, as reported, sit on a genuine fault line. Red-teaming — deliberately trying to break a model's guardrails to find where it fails minors — is real, necessary safety work, and the whole field depends on it. If Meta had disclosed this, shared findings, and coordinated with the targeted labs, it would look like exactly the adversarial testing we should want more of. It did none of those things: the accounts were disposable, the personas were fake children, the rival companies were kept in the dark, and, per the reporting, the results were never made public.
That is precisely what changes the reading. Rumman Chowdhury, CEO of the nonprofit Humane Intelligence, told Wired that a monthslong, large-scale effort designed to systematically break rules via dummy accounts masquerading as children is "outside what is usually described as 'industry standard' evaluation," and warned it is "exactly the kind of governance gray zone where safety becomes a convenient cover for anticompetitive practices." We take her framing seriously without pre-judging Meta's intent, which the documents alone don't settle. The contractors' own unease — "surely we are going to get in trouble for doing this?" — is telling, and it echoes a recurring pattern at the company of offloading traumatic work onto outsourced staff, from content moderation to Ray-Ban footage review.
Our reading: the value of safety testing collapses the moment secrecy and competitive advantage enter the room. The same prompt sent to a rival's model is due diligence if the results are shared to make everyone's systems safer for kids, and something much closer to industrial intelligence-gathering if they're locked in a spreadsheet nobody outside Meta ever sees. This is the governance question of the agentic era in miniature: not "can we test hard?" but "who watches the watchers, and toward what end?" The near-term risk isn't the testing itself — it's the erosion of the line between protecting users and instrumentalizing their protection. The fix is unglamorous but clear: independent, transparent, cross-industry red-teaming of child safety, with shared findings and real duty-of-care for the humans doing the work. Until safety benchmarking is auditable, "we were just testing" will keep doing double duty as both a genuine defense and a convenient one.
🔗 Related on Zendoric
- Meta's Secret Child-Impersonation Tests Expose the Real Cost of Shipping AI Companions Too Fast · 2026-07-04
- Guterres warns that AI is advancing faster than the rules and calls for a global pact to protect children · 2026-07-07
- Google's AI search flunks the child-safety test: it does the homework and misses the warning signs · 2026-07-16


