7. TrustLayer in Chat vs. Automated Workflows
Distinguish between annotation-only mode in chat and agents, and blocking mode in automated workflows where a control can stop unverified output before delivery.
TrustLayer operates in two distinct modes depending on where the AI is working.
Annotation-only mode (chat & agents): When you interact with AI in the chat or through agents, TrustLayer scores and annotates every answer — but it never blocks or modifies the output. You see the badge (green, orange, or red), you read the details, and you decide what to do. The system trusts you, the human in the loop, to act on warnings.
Blocking mode (automated workflows): When AI runs unattended — scheduled tasks, automated pipelines — there is no human watching each output. Here, a TrustLayer control can actually block delivery of an unverified answer before it reaches its destination. This prevents a workflow from publishing or sending content that carries a red-level issue without anyone reviewing it.
The key distinction: in chat you are the safety net; in workflows, TrustLayer becomes the safety net.
Why does the mode matter? Consider the five controls TrustLayer runs — source anchoring, prompt fidelity, action claims, code guard, and repetition. In chat, if source anchoring flags an ungrounded claim (orange badge), you can immediately ask the AI to cite sources or attach a document. You adapt in real time.
In an automated workflow, nobody is there to react. If the action-claims control detects the AI said it sent an email but no send actually occurred, or if code guard spots a dangerous pattern in generated code, the blocking mechanism stops that output from going live. Without this, a scheduled report could silently distribute fabricated data or unsafe code.
In chat, treat an orange badge as a prompt to investigate — click the badge, review which control flagged the issue, and ask the AI to rephrase or cite sources. Treat a red badge as a stop sign: never forward a red-badge answer without manual verification. In workflows, the platform enforces this for you automatically.
Remember that short conversational replies (e.g., "sure, which one?") contain no verifiable claim. TrustLayer marks these as not applicable rather than assigning a misleading score. This applies in both modes — annotation and blocking — because there is simply nothing to check. The badge stays meaningful by not scoring the trivially empty.
A green badge does not mean "true in the world." It means the answer is consistent with the sources, request, and tools TrustLayer can see. Even in blocking mode, a green-pass output could contain a factual error if the underlying source was wrong. Your judgment on high-stakes decisions remains essential regardless of mode.
Open a chat conversation and generate an answer grounded in a document or knowledge base. Click the TrustLayer badge to inspect each control's result. Then imagine the same answer running in an automated workflow: would it have passed or been blocked? This exercise builds your intuition for when annotation is enough and when blocking is critical.
In chat and agents, TrustLayer annotates only — you see the score and color, and you decide how to act. In automated workflows, TrustLayer can block unverified output before delivery, protecting against unsupervised errors. Short replies are marked not applicable in both modes. Green means internally consistent, not objectively true. Click the badge to see per-control details, and always apply your own judgment on critical outputs.