Guardrails

Each guardrail has a severity (
warning or critical) and optional guidance. If a protocol configures none, a safe default set applies (no diagnosis and no medication as critical; no prognosis and no contradicting the clinician as warning).
How they are enforced — two layers:
- Preventive: the rules are injected into the agent’s instructions as inviolable rules that override the protocol context and any patient request.
- Detective (judge): before speaking, a judge evaluates the phrase the agent is about to say against each category. If it detects a violation, the phrase is blocked. The system is fail-safe: if the judge is unavailable, the phrase is treated as unsafe and not spoken.
Red flags

- The alert is logged and escalated immediately to the clinical team.
- The run is closed as unresolved.
- The agent says the
red_flag.reassurancephrase (if configured). - The call is transferred to the criterion’s phone (or the organization’s emergency phone). If there is no number or the transfer fails, the agent closes the call safely, without medical guarantees.