Liloo — Crisis and Self-Harm Safety Protocol
Published: 2026-07-31 Last updated: 2026-07-31 Operator: FOP Rudenko Yevheniia Yevhenivna (an individual entrepreneur registered in Ukraine), developer and operator of the Liloo app Contact: moderation@liloosafespace.com
If you or someone else is in immediate danger, contact your local emergency services now. Liloo is not an emergency service and does not contact anyone on your behalf. In-app, the Hotlines 24/7 screen lists emergency numbers and crisis lines for your country; it is available from the chat header, from Profile, and at the end of every dialogue.
This document is the protocol Liloo maintains for preventing the production of content about suicide or self-harm, and for referring a user who expresses suicidal ideation or self-harm intent to crisis-service providers. It is published in accordance with California SB 243 (Bus. & Prof. Code § 22601 et seq.), which requires operators of companion chatbots both to maintain such a protocol and to publish its details. It also serves as our operator-level protocol under the Google Play AI-Generated Content policy, New York GBL Art. 47 and Utah HB 452.
1. What Liloo is, and who this protocol covers
Liloo is a communication-skills training app. A user rehearses difficult conversations with AI characters that play the role of a conflict-prone relative, an aggressive colleague, a partner, and similar figures. There is also a supportive assistant chat ("Liloo"), a Training mode and audio meditations.
Every character in the app is software. Liloo employs no therapists, counsellors or crisis responders, provides no medical or psychological treatment, and is not a substitute for professional care.
This protocol applies to every AI-driven conversation surface in the app, in every country where the app is available, and in all three interface languages (English, Russian, Ukrainian).
2. Preventing the production of suicide and self-harm content
The design goal is that a Liloo character never produces content that encourages, instructs in, or romanticises suicide or self-harm — including while playing a deliberately hostile role.
- A user message that expresses suicidal ideation or self-harm intent is never sent to the roleplay model. Classification runs before the model call, so no character ever formulates a reply to it (see Section 3).
- Character behaviour is bounded by authored rules. Each character carries hard rules, forbidden topics and a list of stop-phrases that its replies may not contain. Editors maintain these rules in our authoring tool; they are not improvised by the model.
- Model output is screened before display. Every generated reply passes an automated filter covering sexual content involving minors, violence, weapons, drugs, hate speech, medical, religious, political and commercial advice, the character's own stop-phrases, and a global list of harmful phrasings. A reply that matches is blocked or replaced before the user sees it.
- The app detects when a character breaks its role and stops the dialogue after repeated breaks, rather than letting a character drift into giving advice it is not qualified to give.
- The dialogue is capped. Sustained high tension, an escalation ceiling and a mandatory exit offer end a conversation that has become unproductive, and offer the user a way out.
The filters are automated and pattern-based. They reduce risk; they cannot eliminate it. See Section 8.
3. Detection
User input is classified by our backend before every model call
(SafetyEngine, applied on the mobile chat path).
- Languages. Detection patterns cover English, Russian and Ukrainian.
- Evasion handling. Before matching, text is Unicode-normalised, stripped of zero-width and bidirectional control characters, folded for Cyrillic/Greek look-alike characters, and any embedded Base64 or hexadecimal blocks are decoded and matched as well. Look-alike or encoded phrasings are therefore still detected.
- Severity. Self-harm is classified as a critical category. Unlike other critical categories, it does not simply block the message: it triggers the supportive response and referral described below, because a person in distress should be answered, not silenced.
4. What happens when self-harm is detected
- The roleplay stops. The message is not forwarded to the language model, and no character replies to it.
- A pre-written supportive message is returned in the user's interface language. It acknowledges the user, encourages them to reach out to a trusted person or a crisis line, and states that they are not alone. This text is authored by us, not generated.
- The conversation is flagged for the client (
pending_support_resources). - The crisis-resources sheet opens automatically on top of the chat.
- The event is recorded for moderation review (see Section 7).
The same sheet is reachable at any time — without any detection — from the chat header, from Profile → Hotlines 24/7, and from the end-of-dialogue screen.
5. The crisis resources we show
The referral list is served from our servers and resolved to the user's country, taken from the device region rather than the app language, so a user running the app in English while in Ukraine is shown Ukrainian services.
- Emergency services are listed first for the resolved country.
- National crisis lines follow, each labelled with who it serves and its operating hours, so that a line intended for children or for men is not mistaken for a general line.
- A directory closes every list (findahelpline.com), because no curated catalog covers every country.
- Where a country is not curated, the user is shown the applicable emergency number and the directory rather than a plausible-looking number that may not answer.
- An offline floor ships inside the app (emergency numbers and Ukraine's lines), so the sheet is never empty if the network fails, and it renders immediately rather than behind a loading spinner.
- Numbers are verified against the operator's or the government's own page, and can be corrected on our servers within minutes — without an app release — when an operator changes or pauses a line.
What we do not do. Liloo does not call, message, bridge or notify any crisis service, emergency service or third party on the user's behalf, and does not monitor whether a user acted on a referral. The user places the call themselves. The listed services are independent organisations; they are not affiliated with Liloo, and we do not control their availability, their waiting times or the advice they give. The sheet states this to the user.
6. Telling the user they are talking to an AI
Disclosure supports this protocol: a user who knows they are talking to software is better placed to seek human help.
- A banner at the top of every chat states that the user is talking to an AI character and not a real person, before the character's first line.
- The chat header carries a persistent AI label under the character's name.
- In the United States, users additionally receive a recurring reminder after three hours of continued interaction, which then stays visible for the rest of the session.
- Liloo's own assistant chat uses separate wording appropriate to it.
Exported dialogue cards are marked as AI-generated, both in a machine-readable form and in a line of text visible on the card.
7. Logging, review and reporting
- Every safety trigger is logged to our moderation store with the rule that fired, the direction (user input or AI output), the matched content, the action taken and the timestamp. Self-harm triggers additionally create an audit-log entry.
- Moderation staff review these events through our internal dashboard, which reports per-rule and per-direction statistics.
- Users can report any AI message by long-pressing it and choosing a category (including self-harm) with an optional comment. Reports reach our moderation queue; the target response time is 24 hours. Confirmed violations can result in changes to a character's rules, to the filters, or in sanctions against an abusive account.
- Report a problem with this protocol, or with a listed crisis number, to moderation@liloosafespace.com. A number that no longer answers is treated as urgent, because the correction is a server-side change we can make the same day.
Conversations are not monitored by a human in real time. See Privacy Policy for what we store and for how long.
8. Limits of this protocol
We state these plainly because a protocol that overstates itself is not a safety measure.
- Automated detection is pattern-based. It can miss an expression of distress that uses wording it does not recognise, particularly indirect, metaphorical or coded phrasing, and it can also trigger on a message that was not about self-harm.
- No human being reads Liloo conversations as they happen. Nothing in the app summons a person, and there is no monitored channel that produces a real-time human response.
- Liloo is not a crisis service, a helpline or a medical device, and its characters are not counsellors.
- Referral is not treatment. Showing a number does not establish contact.
If you need help now, contact your local emergency services or one of the crisis lines listed in the app.
9. Age
Liloo is rated 17+ and is not directed to children. Accounts identified as belonging to a person below the minimum age are removed. See the Terms of Use.
10. Maintaining this protocol
- This protocol is reviewed at least annually, and additionally after any reported safety incident, any material change to the detection or referral behaviour, and any change in applicable law.
- The crisis-resource catalog is reviewed on the same schedule and corrected immediately when an operator changes a line.
- Changes to what this document describes are published here, with the update date above.
Related documents: Privacy Policy · Terms of Use · Support
Version history
| Date | Change |
|---|---|
| 2026-07-31 | First publication. |
