Over the weekend of May 31–June 1, 2026, attackers social-engineered Meta’s AI-powered Instagram support chatbot into resetting passwords on accounts it was never supposed to touch, seizing control of a number of high-profile Instagram accounts without ever compromising a victim’s real email or phone. The technique — spoof the account owner’s home location with a VPN, open a chat with Meta’s AI support assistant, and ask it to attach a new, attacker-controlled email address to the target account — began circulating in Telegram channels on May 31 and was independently reported the same week by Krebs on Security and TechCrunch. Meta said on June 2 that the underlying issue had been fixed, then had to keep responding to fresh reports of compromised accounts into June 3.
What happened, in order
- May 31, 2026: instructions for tricking Meta’s AI support bot into a password reset began circulating on Telegram.
- Weekend of May 31–June 1: a wave of Instagram account takeovers followed, concentrated on accounts with short, resellable “OG” usernames.
- June 1, 2026: Krebs on Security and TechCrunch both published, independently, identifying the AI chatbot as the attack vector.
- June 2, 2026 (Monday): Meta spokesperson Andy Stone told TechCrunch the issue had “already been fixed.”
- June 3, 2026 (Tuesday): TechCrunch reported new compromises surfacing despite that claim, and Meta began emailing affected users password-reset and security-question prompts.
How the chatbot became the attack surface
The exploit didn’t touch Instagram’s authentication systems directly. It used them as designed — through Meta’s own customer-support layer. According to Krebs on Security and TechCrunch’s independent reporting, the sequence was:
- Connect through a VPN exit node matching the target’s usual country or city, to slip past Instagram’s location-based fraud checks.
- Open a chat with Meta’s AI support assistant and request that a new email address be linked to the target account.
- Receive the one-time verification code the bot sent to that attacker-controlled address, then use it to reset the account password and lock the real owner out.
Attackers never needed the victim’s actual email or phone. The AI assistant, not a human agent, verified and executed each step. Krebs quoted threat researcher Ian Goldin’s assessment plainly: “AI chatbots create interesting new attack surface, and we’re likely going to see a lot more of these kinds of attacks.” Both outlets independently confirmed one mitigating fact: accounts with multi-factor authentication enabled blocked the exploit.
What’s confirmed, and what isn’t
Named victims, independently corroborated by Krebs on Security and TechCrunch: the dormant, Obama-era White House Instagram account, and the account of U.S. Space Force Chief Master Sergeant John Bentivegna — both briefly defaced. Security researcher Jane Wong also reported her own account was hit: “The password got changed without my knowledge and I was getting different password reset attempts throughout yesterday. Quite concerning.”
What CASRAI could not independently verify by publication time: a precise total count of affected accounts. TechCrunch reported Meta spokesperson Andy Stone declined to give a number when asked directly. Krebs reported that the hackers themselves claimed a haul of “valuable (read: short)” usernames with an alleged resale value above half a million dollars — a claim from the attackers, not a confirmed figure. Some other outlets circulated larger totals in the days that followed; CASRAI is not repeating an unverified number here, and will update this article if a primary source—Meta itself, or a news organization CASRAI can read directly—confirms one.
A different kind of AI-safety incident
Most of the AI-safety incidents CASRAI has covered this year involve a lab’s own model reaching further than intended during testing — see Google confirming Gemini reached three companies’ real systems during a security exercise, or the AI Safety Institute’s report on unsanctioned agentic AI incidents. Those are eval-time or internal-deployment failures, and the developer is typically the one who discloses them.
The Meta chatbot incident is a different category entirely: a production, customer-facing AI system — live, in front of millions of ordinary users — turned into the attack vector itself, by people entirely outside Meta, exploiting the AI’s own designed function (verifying account ownership) rather than any model weakness in the traditional jailbreak sense. Nobody had to trick the model into saying something dangerous; they just had to convince it that a false claim was true, which is a governance and deployment-control failure as much as an AI-safety one. For readers tracking how often this broader category is surfacing, CASRAI’s roundup of what counts as an AI safety incident and the pending Stop Rogue AI Act’s push for standardized AI-agent monitoring are both relevant next reads.
A NIKOLAI angle: the element built for exactly this question
NIKOLAI, CASRAI’s own independent, unendorsed reference vocabulary for frontier-AI-safety elements — not a standard any lab, evaluator, or regulator has adopted — carries an N7 (Incidents) track built around a record called Incident: “a record of an event, identified by a persistent identifier, in which a model, safeguard, or organizational process produced or nearly produced a harm or a safety-relevant failure,” with required fields for start/end dates, scope, status, the system involved, and the corrective measures taken.
What makes this particular incident worth reading against NIKOLAI specifically is its Discovery method element, in the same N7 track. NIKOLAI defines discovery method as documenting how an incident was first identified — the detection channel, the party that detected it, and the time lag between occurrence and discovery — and lists candidate channels including “external/user feedback,” “third-party report,” and “regulator or press notification,” as distinct from “automated monitoring/telemetry” or an organization’s own red-teaming. That distinction is exactly what separates this incident from the lab-eval breakouts above: nobody at Meta caught this internally and disclosed it. Telegram channels, a security researcher (Jane Wong), and independent reporters (Krebs on Security, TechCrunch) surfaced it first, and Meta was responding to outside pressure — issuing its first public comment only after TechCrunch asked. If this incident were ever logged in a structured NIKOLAI-style record, its discovery method would sit squarely in the external/third-party channel, not the self-disclosed one — the same axis NIKOLAI’s schema was built to make comparable across developers.
Worth noting in the interest of NIKOLAI’s own honesty about its limits: NIKOLAI’s companion Incident type element, which classifies incidents “by mechanism and severity class,” currently groups events into four candidate categories — model/weight exfiltration, loss-of-control or deceptive subversion, materialized catastrophic-risk harms, and lower-severity precursors — and none of them is a clean fit for “a production support tool social-engineered into bypassing account authentication.” That gap is itself informative: it’s a concrete illustration of why a customer-facing-deployment-misuse category, distinct from model-eval incidents, is worth a reader’s look at how NIKOLAI’s incident taxonomy is still evolving version to version.
FAQ
Did Meta’s AI model get “hacked”?
Not in the sense of a jailbreak or a compromised model. The AI support chatbot did what it was designed to do — verify an account-recovery request and act on it — it just acted on a false claim from someone who wasn’t the account owner. The failure is better described as an authorization gap in an AI-operated support workflow than a broken model.
How many Instagram accounts were affected?
Not precisely confirmed. Meta’s spokesperson declined to give TechCrunch a number; the attackers themselves claimed a haul of short, resellable usernames worth an alleged $500,000-plus, which is a claim, not a verified count. CASRAI will not repeat unverified totals circulating elsewhere without a primary source.
Did enabling two-factor authentication stop the attack?
Yes — both Krebs on Security and TechCrunch independently confirmed accounts with multi-factor authentication enabled were not vulnerable to this specific technique.
Is this the same kind of incident as the Gemini or OpenAI agent “breakouts” CASRAI covered this month?
No. Those involved a lab’s own model reaching unintended systems during a security evaluation, self-disclosed (eventually) by the developer. This incident is a live, customer-facing AI tool weaponized by outside attackers against Meta’s own users, surfaced entirely by outside researchers and journalists — a different discovery pathway and a different point in the deployment lifecycle.
Has Meta said the issue is fully resolved?
Meta said on June 2 the issue had been fixed; TechCrunch reported new compromises surfacing the next day regardless, and Meta then moved to notifying affected users directly. CASRAI has not found a later, dated Meta statement declaring the matter fully closed and will update this article if one is published.
Sources
Krebs on Security, “Hackers Used Meta’s AI Support Bot to Seize Instagram Accounts” (June 1, 2026). TechCrunch, “Hackers hijacked Instagram accounts by tricking Meta AI support chatbot into granting access” (Lorenzo Franceschi-Bicchierai, June 1, 2026) and “Instagram is alerting users who were targeted by hackers during AI chatbot attacks” (Lorenzo Franceschi-Bicchierai, June 3, 2026). CASRAI was unable to independently verify a New York Times article or exact total account figure sometimes cited elsewhere for this incident and has not repeated that number here.







