Safety Center

Responsible AI for Parents

Automated safety tools, including ours, misread language
constantly.

They struggle with sarcasm, inside jokes, song lyrics, gaming trash
talk, reclaimed slurs used between friends, regional dialect, and
ordinary sibling conflict. A model that never misses a real threat will
also flag a hundred harmless messages.

That trade-off is the reason confidence levels exist. A
low-confidence flag deserves a glance. A high-confidence pattern from an
unknown account deserves your full attention.

How to use these tools well:

  • Treat an alert as a prompt to look, never as proof
  • Read the actual context before reacting
  • Never punish based on a flag alone
  • Tell your child that alerts get reviewed by you, not acted on
    automatically
  • Mark false positives so you learn the tool’s blind spots

The failure mode to avoid is treating the dashboard as a verdict
machine. If your child learns that a misread joke produces a
consequence, they will move the conversation somewhere you cannot see,
and you will have traded real visibility for the appearance of it.

Start the conversation: “The app flagged this and it may have read it wrong. What was actually going on?”

This article is educational. It is not medical, legal, or
emergency advice.


← Back to Safety Center