Rules an AI agent cannot break: why a prompt is a suggestion, and how a gate before every tool call and every send makes a rule real
A rule in a prompt holds most of the time. A rule in code — a gate before each tool call, a gate before each send, a verifier on each reply — holds every time. What that looks like, the rule effect that does nothing, and why the gate fails closed.
A rule written into an AI agent's prompt is a suggestion the model usually follows. A rule the agent cannot break is enforced outside the model: a gate that checks every tool call before it runs, a gate that checks every outbound message before it sends, and a verifier that checks every reply against what the agent actually retrieved. This is how a business can let an AI talk to customers without reading every message first.
On this page
The problem with "just tell it"
"Never quote a price not in the catalogue." "Always hand over when someone mentions a lawyer." "Do not promise a delivery date." Put those in a prompt and the agent will follow them in nearly every conversation — and the one it does not is the one that reaches a customer, because nothing sits between the model's output and the send button. A prompt is a strong hint. It is not a rule, because nothing checks it.
Three places a rule can actually be enforced
- Before a tool runs. The agent proposes an action — look up a product, write a field, request a payment, advance a stage. A gate in code checks the proposal against the rules: is this tool allowed for this conversation, does this field exist, is this contact blocked, does this topic force a handover. A refused call is refused visibly, so the model knows it did not happen.
- Before a message sends. A second gate checks the outbound message: does the person consent to this kind of message on this channel, are they suppressed, is the 24-hour window open, is the account allowed to send at all. Every message from anywhere — the agent, a campaign, a reminder — takes this path; there is no other.
- Before a reply is accepted. A verifier reads the reply against what the agent retrieved this turn. A price the agent did not look up is not allowed in the reply. A claim about stock, a date, a policy — the same.
What a rule looks like when it is code
| The rule, in words | How it is enforced |
|---|---|
| Quote only catalogue prices | The verifier refuses a reply containing a price unless the turn retrieved that product; a payment request is priced from the catalogue, not from the chat |
| Hand over on hazardous cargo / a legal threat / a complaint | A topic rule sets the conversation to human before the agent replies; the agent's next turn is skipped |
| Never message someone who said STOP | Suppression is checked at the send gate, for every sender, every time |
| Do not shortlist before budget is known | The shortlist tool is blocked until the budget field is filled; the agent is told which field is missing |
| Anything that spends money needs a person | A campaign, a rule change, a catalogue edit proposed by the agent becomes a pending approval, never a change |
| A closed set of tools | The agent has five tools and cannot be given a sixth by configuration; a new capability is a code change with an argument |
The rule effect that does nothing
The failure to watch for is a rule with no input. A "hand over on legal topics" rule only works if something classifies topics and writes the result where the gate reads it. A "needs approval" effect only works if approvals are written somewhere a person actually looks. In a product this was built from, all three existed and nothing wrote to them — escalation did not stop the agent, approvals reached nobody. Wiring a rule means checking what produces its input, not just that the branch exists.
Fail closed
Every rule effect has a branch, and the gate's fallthrough refuses. An effect nobody implemented is a blocked tool, not an ungoverned one. That inverts the usual failure: a bug in the rules layer produces an agent that does less, never one that does more.
What this buys a business
- You can write a rule in a sentence on the dashboard and know it holds on every turn, not most.
- You can let the agent run overnight without reading the transcripts in the morning, because the things that would have been wrong were refused rather than sent.
- You can see why something did not happen: a refused call names the rule.
- The agent cannot widen its own permissions. It may draft a rule that would let it do more; a person says yes.
Can I add my own rules to the AI?
Yes, in words, on the dashboard: which topics hand over, which fields must be filled before which action, what the agent may and may not do for which kind of contact. Each becomes a condition the gate evaluates in code on every turn.
What happens when a rule blocks the agent?
The tool call is refused and the agent is told why, so it can ask for the missing field or hand over. Nothing is silently dropped — a call the model believes ran and did not is worse than one it knows was refused.
Does this make the AI worse at conversation?
It makes it worse at the conversations you did not want — quoting a discount, promising a date, arguing with a complaint. In the ones you did, it asks better questions because it knows what it is not allowed to guess.
Read next
More on Selling with an AILinkubit vs other toolsWhatsApp templatesCost calculatorPricing
Try it on your own number. Linkubit answers, qualifies and quotes on WhatsApp from your catalogue, and adds nothing to Meta’s per-message price. Start free — fourteen days of Pro, no card — or read the pricing.