What happens when a rehearsal gets difficult

Liloo lets you practise conversations that are hard to have: with a partner, a parent, a manager. The character on the other side is allowed to push back, because a rehearsal where nobody pushes back teaches nothing. This page explains what runs underneath that, and where it stops.

It is written for anyone deciding whether to trust the app with a real problem — and it is deliberately specific, because "your safety is our priority" tells you nothing.

First, what Liloo is not

Liloo is a communication-skills trainer. It is not therapy, not counselling and not medical care. It does not diagnose anything, it does not treat anything, and it is not a substitute for a professional. If you are in danger right now, an app is the wrong tool — the crisis lines below exist for that.

Every message passes two filters

Before the AI sees it. Your message is normalised first — unicode folded, invisible characters stripped, look-alike letters resolved, encoded payloads decoded — so that a filter cannot be evaded by writing the same thing in a different alphabet. The normalised text is then checked for self-harm signals, sexual content involving minors, incitement to violence, and attempts to talk the character out of its role. Some of these stop the exchange outright. Others change what happens next, which is the section below.

Before you see it. The character's reply passes the same normalisation and its own set of checks: topics the persona must not raise, phrases that are known to re-traumatise in that specific role, and replies where the character has fallen out of its role and started explaining that it is a language model. Two consecutive breaks of that kind stop the scene and recover the role rather than letting it drift.

Neither filter is a language model judging a language model. They are explicit, inspectable rules, which means they are predictable and can be audited — and also that they will never catch everything.

When the conversation signals distress

Three responses run in production, and they are different from each other:

  • Self-harm or crisis signals route out of the roleplay entirely and surface support resources for your country, not a generic list. The full protocol is published separately: see the safety protocol page.
  • A run of aggressive turns — three in a row — offers a support line rather than continuing to escalate.
  • High escalation inside a scenario triggers a mandatory exit from the scene, so a rehearsal cannot spiral indefinitely.

You are always told it is an AI

The character is a simulation, and the app says so: in the chat itself, in the header, and — for users in jurisdictions that require it — as a periodic reminder. No part of the product tries to make you forget what you are talking to, and the character is not a digital copy of any real person, including the person you have in mind.

When something goes wrong anyway

Every AI message has a report control. A report goes to a human moderation queue with the message attached. This is not decoration: it is the path by which the filters above get extended, and reports are the main source of new patterns.

The limits, stated plainly

  • The filters are rule-based. Rules match what they were written to match; novel phrasings get through until someone adds them.
  • The hate-speech catalog currently runs a baseline of unambiguous incitement patterns, not a comprehensive slur list.
  • A generative model can produce a sentence nobody anticipated. That is a property of the technology, not a bug we have left open.
  • Nothing here constitutes clinical judgement. The system detects patterns in text; it does not understand your situation.

We would rather write that down than claim a completeness we cannot demonstrate.

If you need help now

The app shows crisis resources for your country. If you are reading this on the web and you need help immediately, contact your local emergency number or a national crisis line — and please reach out to a person, not to a rehearsal.