Understanding the activity feed
The Activity page is your window into how the safety rules are working. It shows safety events, moments when a rule looked at something, never the conversations themselves.
The events
Each entry shows the time, which family member, which AI (ChatGPT or Claude), and what happened:
- Blocked: a rule stopped a request or an AI answer. Your family member saw a notice with your rule's message (and your name, if you've set one on the Members page). The AI never saw the blocked request.
- Waiting for approval / Approved / Denied: an ask-first rule held a request for your decision. The entry links to the outcome, so you can see how it ended.
- Noted: an allow-and-notify rule let something through but recorded it for you.
Entries show the rule that acted (like “Block Weapons”) and the matched term where one exists. They never show the conversation.
“Possible:” chips, the amber ones
Sometimes you'll see an amber chip like Possible: Block Violence. This means the AI classifier looked at a message and thought it might involve that topic, but wasn't confident enough for the rule to act. Concretely: each intent rule has a sensitivity threshold (75% by default). Matches above it act. Matches below it, but above a noise floor, become “Possible” chips:
- Nothing was blocked; the message went through normally
- Your family member saw nothing
- No alert was sent
Think of it as the classifier's “I wasn't sure” pile. It exists so you get visibility into borderline moments without anyone being punished for the classifier's uncertainty. What to do with them: mostly nothing. But they're your tuning signal. Seeing “Possible” chips on things that should have been blocked? Lower that rule's sensitivity threshold on the Policies page. Seeing them on obviously innocent things (video-game talk is the classic)? The threshold is doing exactly its job; leave it.
Shield badges
When a rule acts on a message, the member's chat marks that message with a small shield. It's the honest-by-design marker: your family member always knows when a rule stepped in, and the same visual connects the message to the notice they saw.
Trend alerts
Some rules watch for patterns instead of single messages, like repeated profanity. When a rule's threshold is crossed (say, 5 matches in 30 minutes), you get one trend alert instead of five pings. These appear in the feed like other events.
What the feed never shows
Conversations. The feed is built from safety-event records: labels, times, rule names. The one place you'll ever see message content is the short preview on an approval request, because you can't decide on what you can't read. Everything else lives only on your family's own devices.
