Skip to content

Power

Your Chatbot Can Flag You. The Rules Are Mostly Hidden.

OpenAI has described a pathway from threatening chat to human review and possible police referral. Most rivals reserve similar powers without explaining the threshold.

Kurt HalloranPower — Politics & Media

August 11, 2026 · 8 min read

A laptop showing a chatbot privacy policy beside a phone displaying an emergency disclosure page.
A laptop showing a chatbot privacy policy beside a phone displaying an emergency disclosure page.

The most important line in OpenAI’s public safety material is easy to miss. Conversations that appear to involve a serious threat against another person may be routed to a specialized human review team, which can suspend an account and refer an imminent threat to law enforcement.

That sentence is the hinge. It turns a chatbot from a responsive product into an intake desk for a private threat-assessment system, one whose evidence comes from intimate conversation and whose rules were written by the company operating it. The user may think they are talking to software. The software company has reserved the right to become a witness.

This is often described as a safety feature, which is not false. A credible plan to hurt someone should not be treated like a request for a pasta recipe. But the safety promise leaves the difficult part off the label: a model or moderator must interpret tone, fantasy, venting, coercion, role-play, delusion and intent, then decide whether the police are the least dangerous next step.

The report begins before any police officer appears

A chatbot does not usually dial emergency services because one alarming sentence crossed the screen. The visible response, such as a refusal or crisis-resource message, is separate from the platform’s internal response.

First comes automated detection. A classifier, meaning a model trained to sort content into risk categories, may score the conversation for self-harm, violence, exploitation or another prohibited use. A sufficiently high score can trigger friction inside the chat, limit an answer, create an internal flag or send material to a human reviewer. Providers rarely publish the thresholds, false-positive rates or exact amount of surrounding conversation shown to that reviewer, partly because detailed rules would help malicious users evade them.

That secrecy also keeps ordinary users from knowing what evidence the company is assessing. One message can look different beside the preceding hour of conversation. Account history, prior warnings and technical signals may affect an enforcement decision, depending on the service. The reviewer can then close the flag, restrict the account, preserve records or escalate the case.

Preservation matters. It means retaining existing data so it remains available if valid legal process arrives later; it is not the same as handing everything to police immediately. The pipeline has several gates, but every gate belongs to the platform.

Return to that OpenAI sentence. A specialized review team sounds reassuring because specialization suggests care, yet the public description does not establish a clinical credential, an evidentiary burden or an appeal available before an emergency referral. It describes an internal corporate function making a high-stakes judgment from text.

OpenAI says the quiet part more clearly

Among major consumer chatbot companies, OpenAI has offered one of the clearest public descriptions of proactive escalation. Its published approach distinguishes threats against other people from self-harm: serious threats to others can receive human review and possible law-enforcement referral when reviewers determine the danger is imminent, while the company has said it does not currently refer self-harm cases to law enforcement.

That distinction matters. Police intervention during a mental-health crisis can introduce detention, force and criminalization into a situation that needs care. OpenAI’s stated restraint on self-harm referrals recognizes part of that risk, although words such as serious and imminent still depend on internal interpretation when violence toward others is involved.

The company’s usage policies separately allow warnings, restrictions and account termination. Its privacy policy also permits disclosure when required by law and in certain emergencies involving danger. These are different authorities: moderation rules govern whether someone may use the product, while privacy terms and law govern when stored information may leave the company.

No public policy can show how consistently workers apply it. Nor do broad transparency figures usually reveal how many emergency disclosures began with chatbot content, how many involved mistaken interpretation or how often a referral produced no finding of danger. The sentence tells users that the pathway exists. It does not provide an audit of the pathway.

The other chatbots reserve power without mapping it

Anthropic, Google, Microsoft and Meta all publish rules against violent or dangerous misuse, and their privacy materials generally reserve the ability to disclose information to comply with legal demands or protect people from serious harm. That does not mean each company operates the same proactive reporting system OpenAI has described.

Anthropic explains that Claude conversations may be subject to automated safety systems and review under its usage rules, while its privacy terms allow disclosures connected to safety, legal obligations and protection from serious harm. Its public material does not offer consumers an equally concrete map from a threatening Claude exchange to an unsolicited police referral.

Google tells Gemini users that some conversations can receive human review to improve services and enforce policies, warning users not to enter confidential material they would not want reviewed. Google’s broader privacy framework allows disclosures for legal process and protection against harm. Microsoft applies safety systems and service rules to Copilot, then relies on its privacy statement and law-enforcement request procedures for disclosures. Enterprise versions can carry different data protections from consumer products, an important distinction when the same product name appears on both.

Meta’s AI products sit inside a larger account and advertising apparatus. Its terms and privacy policies allow content review, account enforcement and disclosures connected to legal requests or safety, but the public-facing material does not give users a detailed threat-escalation flowchart. Character.AI, whose conversational products have drawn particular scrutiny around young and vulnerable users, likewise combines automated moderation, account enforcement and privacy-policy emergency clauses without providing a public, case-level ledger of escalations.

The pattern is not that every chatbot secretly calls police. The pattern is contractual capacity. Companies collect conversations, screen them, keep some of them, let workers inspect selected exchanges and preserve broad discretion to disclose data when their lawyers or safety teams believe an emergency standard has been met.

The legal standard is permission, not a duty of care

In the United States, the Stored Communications Act regulates when service providers may disclose stored communications and customer records. Its emergency exception permits a provider to disclose information voluntarily when it believes in good faith that an emergency involving danger of death or serious physical injury requires disclosure without delay.

Permits is doing substantial work. The exception is not a universal command to report, and a chatbot company is not automatically operating under the professional duties that can apply to licensed clinicians. A user also should not assume that a conversation with a general chatbot carries therapist-patient privilege. The interface may imitate attention.

The legal relationship does not follow the performance.

Police can also request data through ordinary legal process. The required instrument can vary with the type of information sought: subscriber details, transaction records and message content do not receive identical treatment, while warrants generally carry a higher threshold than subpoenas. Agencies may submit emergency requests without first obtaining the usual process, leaving the provider to decide whether the stated circumstances satisfy its policy and the statute. A later preservation request can stop existing records from disappearing while investigators seek further authority.

This gives two institutions interpretive power. Police describe the emergency. The platform decides whether to accept that description and what data to provide. Users usually see neither exchange in real time, and delayed notice or legally barred notice can make the disclosure invisible for months or permanently.

Safety has an asymmetric incentive

The commercial incentive is plain. A missed credible threat can become a public catastrophe attached to the chatbot’s name, complete with lawsuits, hearings and screenshots showing that the system saw the warning. An unnecessary internal flag is quieter. A mistaken emergency disclosure may impose severe costs on the user, but many of those costs occur outside the company’s dashboard.

That asymmetry encourages expansive detection followed by confidential review. It also explains why companies advertise crisis sensitivity while withholding operational detail: they want credit for intervention, room to change thresholds and protection against people testing the boundary. The result resembles content moderation with a police exit, except the source material can include confessions, intrusive thoughts and emotionally dependent exchanges that users would never post publicly.

A defensible system would disclose more without publishing an evasion manual. Providers could report aggregate emergency referrals by category, distinguish proactive referrals from police-initiated requests, audit disparities, set short retention periods for unsubstantiated flags and offer notice after the danger passes unless law forbids it. Independent reviewers could examine false positives under confidentiality.

The OpenAI sentence remains the useful anchor because it names the transfer point. A conversation becomes a case when a company team decides that private language meets its internal version of imminence. Law supplies an emergency door. The platform chooses when to open it.

Questions people ask

Can ChatGPT contact the police about a user?

OpenAI says a serious threat against another person may be sent for specialized human review and referred to law enforcement if reviewers judge the threat imminent. That does not mean every violent statement triggers a report, and the company does not publish the operational threshold or case-by-case outcomes.

Do chatbots report self-harm conversations?

Policies differ, and crisis messages may trigger supportive responses or human safety review. OpenAI has said it does not currently refer self-harm cases to law enforcement. Other providers commonly reserve emergency disclosure rights, but their public documents often do not explain whether or when a self-harm conversation would produce an unsolicited report.

Can police read chatbot conversations without a warrant?

Police may seek different records through subpoenas, court orders or warrants, depending on the data and jurisdiction. In an emergency involving danger of death or serious physical injury, US law can permit a provider to disclose information voluntarily without waiting for ordinary legal process, if the provider accepts the emergency basis in good faith.

Will a user know that their chatbot conversation was disclosed?

Not necessarily. Notice can be delayed, prohibited or absent during an emergency disclosure, and public transparency reports rarely identify individual chatbot cases. The concrete evidence may be a preserved conversation, account information and technical records selected by the provider, all produced through a process the user never sees.

Was this worth your time?
ShareFacebook
surveillanceinternet policymental healthai chatbotsmental healthprivacylaw enforcement

One update a day

Today's story, in your inbox

One story each morning — no hype, no filler, no algorithm deciding for you.

Read next

A laptop showing an AI meeting transcript beside a calendar invite labeled Weekly Check-In.

Power

Your AI Meeting Notes Can Become Workplace Evidence

A convenience bot can turn one meeting into audio, transcript, summary and action items spread across several systems. Deleting the bot from the call does not delete that second room.

Lena Vasquez · 7 min read

A square pop-up canopy on a campus lawn with one fabric sidewall attached and folded blankets visible underneath.

Power

Campus Protest Rules Now Police the Tent’s Sidewalls

Public universities are recoding protest as a problem of structures, sound and sleeping. The rules look neutral because they describe equipment, while discretion decides whose equipment becomes an offense.

Lena Vasquez · 8 min read