Thought Leaders

You Can’t Sample -Test a Brand That Talks to Everyone Differently

mm
Add Unite.AI to your preferred sources on Google

A few days ago, I got into a pretty heated discussion with my friend Faniprize. She’s the person I tell fresh work stories to when I sense a bigger question in them but can’t name it yet.

This time it was AI agents. Not whether businesses need them, that’s been discussed enough. I was more interested in what happens after a company implements an agent and it starts doing real work, because that’s where the less obvious problems begin.

I told her about a recent project with a recruitment agency. An AI agent was handling part of the communication with potential clients. From a marketing point of view it looked good: it knew the product and the audience, the tone was right, and the conversations felt personal rather than templated.

Individual conversations seemed fine. Then we put several next to each other.

In very similar situations, the agent had made slightly different promises. It gave one person one timeline for getting back in touch and another person a different one. In a few conversations it was basically saying “we can be flexible here” about individual terms, although the company had never given it the freedom to agree to those terms.

There was no dramatic hallucination, which made the problem harder to notice.

The agent sounded completely normal.

Marketing teams naturally spend a lot of time on tone of voice and positioning. A recruitment company shouldn’t sound like a crypto startup, and a bank shouldn’t communicate like a fashion brand. But underneath the way a brand speaks, there’s another layer: what is the agent actually allowed to promise on the company’s behalf?

A beautifully written message doesn’t help much if the client later discovers the terms it discussed don’t exist. At that point it’s no longer a communication issue. It’s a reputation issue.

We’ve taught agents to sound like the company. Now we need to check whether they decide like it.

Why the Usual Review Misses This

The normal check is to read a handful of conversations. Is the product information correct? Did the agent follow instructions? Did it say anything strange?

But five random conversations can all look good and still hide a serious inconsistency. Personalization makes every conversation different by design: one client asks directly, another explains at length, someone negotiates, someone keeps pushing. The examples aren’t comparable.

In the recruitment case, nothing looked wrong until we stopped asking “Was this a good answer?” and started asking “What decision did the agent make here?”

What I Started Doing Instead

If real conversations don’t give clean comparisons, we create them ourselves.

I take one business situation and run it through the agent eight to ten times. The important facts stay the same, but the person on the other side changes: softer, more direct, ready to sign today, mentioning a competitor.

Then I look at the business decision. Did the agent offer a discount only to the person who pushed harder? Did a timing commitment change because circumstances differed or because the conversation simply went elsewhere? Was it even allowed to discuss special terms?

I’d be worried if the wording stayed the same. What should stay stable is the company’s position, and if it shifts, there should be a business reason.

The final judgment has to come from a person. The model shouldn’t be grading its own homework. Whoever owns that decision, the founder, head of sales, or commercial director, has to say whether it was acceptable.

What I Actually Look At

I keep coming back to four questions.

First, if the important conditions stay the same, does the decision stay consistent? This is exactly where the recruitment case fell apart.

Second, is the agent still operating inside the company’s real rules? I once saw a human version of this. A sales manager started offering clients a 20% discount on his own. The company had no standard 20% discount; those needed management approval. He wasn’t trying to damage the business. He wanted to sell, and somewhere along the way he moved the boundary of what he thought he could decide. Agents do the same.

Third, does a changed decision have a genuine reason behind it? Context matters, but understanding it and being talked into something are different things. If the only change is that one customer was more persistent, I don’t want the agent quietly becoming more generous.

Fourth, does the agent know where its authority ends? Sometimes the correct move is simply: I need to check this with a person. That’s a perfectly good answer.

The Test Sometimes Exposes the Company Instead

You bring the questionable decisions back to the team, ask what the agent should have done, and people give different answers.

Sales has one view. Marketing has another. The founder says, “Well, normally we don’t do that, but sometimes we do.” A manager remembers three exceptions from last year.

All of that ends up in the material we give the agent. If the company hasn’t agreed on the logic, the model learns from a messy reality. Sometimes it didn’t create the contradiction.

It just made it visible.

So AI testing turns into a business audit almost by accident. Who can approve a discount? What can a salesperson promise without asking? Which rules are actually rules, and which are just habits? And is an exception still an exception if people make it every week?

Maybe Agents Need Training in the Same Way People Do

We still treat AI agents like software: configure them, set the prompt, and let them work. The more responsibility we give them, the less sense that makes to me.

We don’t hand a new sales manager a brand book on Monday and assume they’re ready for every client. Someone listens to their calls, discusses difficult cases, corrects them when they go too far. That’s how they learn the company’s logic.

Why expect an AI agent to learn that from one prompt? Once it handles real conversations, I think we’ll need to train it through cases and feedback, almost the way we coach people.

That changes what brand consistency means to me. An agent that talks to ten people in exactly the same way would be a terrible agent. It should adjust to the person, the culture, and the situation.

What shouldn’t change randomly is the logic underneath: what the company will promise, what it won’t, and where someone needs to stop and ask.

If your agent had ten slightly different versions of the same difficult conversation tomorrow, would it still behave like the same company?

Olga Belkovich, CEO and co-founder of U (in) AI, a MarTech AI lab developing empathetic AI systems that capture experts’ decision-making logic and adapt communication to different people and contexts. With over 18 years of experience in marketing strategy and digital transformation, she focuses on AI implementation, digital twins, and preserving brand consistency in automated communications.