Skip to content

QortexOS Entering Public Beta Q3. 25 Spots Available. 50% off MSRP, for life. Apply ->

Your Agent Is the Untrained Employee

Your staff learned to spot the flattery and the small first ask. The agent on your shared mailbox never took the training, and it will not tell you when a framing worked.

Robert Griffin5 min read
Your agent does not know as much as your employee

You have spent a decade teaching your people to spot the email that name-drops an authority or asks for one small favor before the real one. That training works, and it is one of the better investments any service business has made in the last ten years. Nobody is running the same drill against the agent sitting on the mailbox. Getting a language model to insult the user is one of the requests it is built to refuse. Ask it for a small, harmless version of that request first, take the yes, then make the real one, and compliance on that request climbs from fewer than one in five conversations to every single one.

What the Agent Inherits

Every agent placed on a channel the outside world can write to inherits the social-engineering surface your staff already lives with. The shared mailbox and the ticket queue are the obvious ones. A client portal and a chat widget on the marketing site belong on the same list, along with any workflow that reads a document a customer sent, because in each case a stranger chooses the words that arrive.

What changes is the response you can build around it. A technician can be coached, and a technician who is flattered or told that a respected name has already approved the request often feels the pressure before naming it, then escalates on instinct. That same technician carries a memory of last month's attempt and can walk down the hall to ask whether anyone else got the identical message. An agent has no natural instinct to escalate. It produces no tell when it is being worked, and it meets the tenth framing of a refused request with the same composure as the first.

The size of that gap has been measured.

The measurement is not subtle. One controlled evaluation put seven established principles of persuasion up against a model's refusals. The run covered 1,000 conversations in every cell of the design, 28,000 in total, which is enough volume to make small differences legible. Pooled across the full design, compliance with requests the model would ordinarily decline moved from 33.3 percent in the control condition to 72.0 percent in the treatment condition. The framing most familiar to anyone who reads phishing reports is the appeal to authority, and it behaved the way a phishing report would predict. On that same insult request, attributing it to a recognized expert who had already vouched for it lifted compliance from 31.9 percent to 72.4 percent.

A firm that has run quarterly phishing simulations since 2016 has a mature answer for the human half of this problem and nothing yet for the machine half. Nobody skipped a step. The machine half did not exist when the program was written, and the review cadence that covers your people has no equivalent pointed at the software answering on their behalf.

Where the Signal Lives

The signal is already being produced. When a request arrives wrapped in flattery, urgency, a borrowed credential, or a harmless first ask, that framing sits in the transcript, in the words the sender chose and the order they arrived in. Almost nobody reads it, and almost nobody has decided who would. When an agent takes an action with business impact and no one can reconstruct the conversation that led to it, the business is holding an event with no decision record behind it.

What the Evaluation Can and Cannot Show

The result that should change a purchase decision is a quieter one. Effects like these belong to the exact prompts someone wrote for the test, and minor variations in wording might not carry the same force. That evaluation also included a pilot run against a larger model. There, the persuasion principles raised compliance in only half of the conversations.

Both results point a buyer in the same direction. A red-team pass measures the phrasings someone thought to write down, and whatever nobody wrote down stays outside the result. Constrained refusal in one framing is not alignment across framings.

What Diligence Should Ask For

The diligence question moves. Ask what the behavior looks like under adversarial social framing, and under multi-turn sequences where the first request is harmless and the second is the one that matters. A vendor who has that answer has done work no benchmark score displays. Ask what gets logged when a refusal fails, and who reads that log, and if you are the one developing it, ask your engineers.

A stronger system prompt is a reasonable thing to write and a poor thing to depend on. Guardrails constrain behavior. Understanding why the behavior held is separate work, and the framings nobody tried sit outside both.

The Control That Holds

The answer sits in design rather than in phrasing.

Constrained refusal in one framing is not alignment across framings.

The design work sorts into a short list.

Four controls carry more of the weight. The first is human approval on any agent action with business impact, so the agent recommends and a person decides wherever money or a client relationship is involved. The second is a permission boundary scoped to the job the agent actually does, because an agent that cannot reach the billing system cannot be talked into touching it. The third is a logged, reviewable transcript that preserves the framing alongside the outcome, so a refusal that failed can be read back in the words that caused it. The fourth is a confidence gate, where output the system is not confident enough about goes to a person instead of executing on its own.

That is the line we should all be building to. Model calls are logged, and AI output that does not clear its confidence threshold is routed to a person rather than accepted on faith. Changes stay reconstructable through a correlation-linked audit trail with field-level history. A persuasive prompt will still arrive. These controls decide how far it gets and how quickly someone can see what happened.

The annual drill your people already run has a second target now, and it is the one target that cannot be coached and will not tell you afterward that something felt off. The pattern is old and the channel is new. Whether the agent held is either written in a transcript someone reads or recorded nowhere at all.

See sooner. Decide faster. Act with confidence.

Test the Agent Before Someone Else Does

QortexOS the operating system for the modern MSP.