Interviews
Sushil Kumar, CEO of Cyara – Interview Series

Sushil Kumar, CEO of Cyara is an experienced enterprise software executive and entrepreneur with more than 25 years of leadership across artificial intelligence, DevOps, cloud infrastructure, product strategy, and software testing. He joined Cyara as CEO in December 2025, following his role as co-founder and CEO of RelicX.ai, where he built a generative AI-powered, intent-based test automation platform that was acquired by Harness. He subsequently led the integration of RelicX’s technology into Harness and helped shape its AI Test Automation strategy. Earlier in his career, Kumar served as General Manager of DevOps at Broadcom, Senior Vice President of Products at CA Technologies, and spent more than 16 years at Oracle, where he held senior product leadership roles and helped scale major enterprise software businesses. Across these roles, he has focused on building and scaling AI, cloud, DevOps, and automation platforms for large enterprises. His appointment at Cyara is focused on expanding the company’s AI-powered customer experience assurance capabilities and global reach.
Cyara is a customer experience assurance company that helps enterprises test, monitor, and validate customer interactions across voice, digital, messaging, and conversational AI channels. Its Cyara Agentic Platform is designed to address the growing challenges created by AI-driven customer experiences, including testing non-deterministic AI agents, detecting hallucinations and behavioral drift, validating compliance, monitoring production systems, and assessing end-to-end customer journeys. The platform combines AI agent testing, production monitoring, voice and telecom assurance, digital-channel testing, and CX observability, supporting more than 350 million customer journeys annually across a global footprint spanning more than 140 countries. As enterprises deploy increasingly autonomous AI agents into customer-facing workflows, Cyara is positioning its technology as an assurance layer for evaluating whether those systems behave reliably, safely, and consistently before and after deployment.
You’ve spent much of your career building and scaling enterprise software, from Oracle and CA/Broadcom to founding Relicx and now leading Cyara. How has that experience shaped your view that AI agents should be managed less like traditional software and more like members of a workforce?
I have spent most of my career building and scaling enterprise software, and the discipline we built there was a discipline around deterministic systems. You know what the software is supposed to do. You validate it against that expectation. When it breaks, it tells you: an error, a failed transaction, an alert.
AI agents do not work that way. They are non-deterministic, so the same input can take a different path. More importantly, they can act on behalf of the company. They make commitments: refunds, policies, promises. And when one of those is wrong, nothing breaks. A wrong answer sounds exactly like a right one. The transaction succeeds, the dashboard stays green, and the customer walks away with something the company never agreed to.
Once software can make decisions and commitments, and can be wrong without telling you, it needs a different operating model.
That is where the workforce comparison earns its keep. You do not manage an employee by scripting every decision they will make. You give them a role, you establish the authority that comes with it, and you expand that authority as they earn it. An agent behaves the same way under the same structure.
My read is that autonomy is not a deployment decision. It is a series of promotions. An agent earns each one by showing it can do the job, stay inside its authority, and recognize when it needs help.
What does an “HR-like” operating model for AI agents actually look like inside an enterprise, and which elements should companies put in place first?
Start with the job. Every agent should have something close to a job description before it goes near production. What is it there to accomplish, what information is authoritative for it, what customer data can it use, what decisions can it make on its own, and where does its responsibility end. If a company cannot write that down in a paragraph, the agent is not ready for a role. It is ready for a demo.
Four things follow from that role, and the order matters. Evidence before launch, which means proving the agent can do the job under conditions that resemble the real world rather than a controlled test. Oversight while it runs, so you know what the agent actually did and not only whether the system responded. Promotion gates, so more authority is granted when there is evidence to support it and not before. And an owner in the business, not in engineering, who is accountable for what that agent is allowed to do.
Get the order wrong and the rest does not hold. If responsibility is vague, good performance is impossible to prove, and so is failure. The role comes first, and the evidence follows.
If an AI agent is assigned a specific role, how should organizations define its responsibilities, permissions, and boundaries before allowing it to interact with customers or critical systems?
The role says what the agent is for. The permissions say what it can reach. Those are two different conversations, and companies tend to have only the first one.
Be explicit about three things. What systems and data the agent can touch, and in which direction, because reading a customer record and changing one are not the same permission. What it can commit to on its own, which is where the money and the liability sit: a refund, a credit, an exception to policy. And what forces a handoff, both the cases you can name in advance and the signal that the agent has moved outside its competence.
These are not decisions to leave to the technology team. They determine the risk the company is taking. The people who own customer experience and the compliance exposure need a say in where those lines get drawn, and they are usually the last ones asked.
Then you have to prove the agent stays inside them. The goal is not to eliminate every possible mistake. There will be mistakes. The question is whether the agent understands its boundaries, knows when to stop, and can do the job it has been given without creating consequences somewhere else in the customer journey.
You argue that greater autonomy should be earned rather than granted from the outset. What should an AI agent have to demonstrate before an enterprise expands the scope of actions it can take independently?
It is now easy to build an AI agent. The hard part is proving it deserves autonomy.
Before expanding what an agent can do on its own, an enterprise needs evidence that it performs its assigned job consistently and stays within its boundaries. That means how it handles the situations you expect, and also the ones you did not anticipate. An agent can look strong under controlled conditions and behave differently when the context or the surrounding systems change.
A customer might begin with a simple billing question and turn frustrated after a failed payment. The agent has to recognize that shift while it is happening and change course, rather than continuing down the path it was validated on.
Three things should be true before authority expands. The agent does the job under real conditions, not only clean ones. It knows the edge of its own competence and stops there. And someone can produce the evidence for both on demand.
The level of proof has to match the level of autonomy. Small decisions, light evidence. Access to a payment system, or the ability to commit the company to a policy exception, and the bar should be considerably higher.
How should companies continuously evaluate the performance of AI agents once they are deployed, especially when the quality of their decisions cannot be captured by traditional software testing metrics alone?
This is where traditional software thinking breaks down. With deterministic software you test whether something passed or failed. With an AI agent you can get a successful response from the system and still have a failed customer interaction.
So you evaluate the outcome, not the response. Did the agent understand what the customer was trying to accomplish? Did it use the right information? Did it complete the journey? Did it stay inside its boundaries and escalate when it should have?
Basic evaluations, scoring answers against a golden set, are the floor. Every company will have those. The dimensions that decide whether a customer keeps trusting you are the ones underneath: compliance, bias, misuse, and how the agent holds up with real callers, their accents, the background noise, the cheap handset, the interruption halfway through a sentence. In voice this matters more than people expect, because every score sits on a transcript. If the speech layer mishears the question, the agent answers one nobody asked.
The arithmetic is worth sitting with. A 99% score in evaluation sounds excellent. At a million conversations a year, that is ten thousand failed ones.
Two principles hold up. The validation should be independent of the agent and model platforms. We do not build the agents ourselves, which is part of why I can say plainly that no vendor should be the judge of its own AI. The standard is the enterprise’s own policies, its customer commitments and its regulatory obligations, not a vendor’s scorecard.
And every production failure should become a gate. Not a ticket, not a backlog item. A test the agent has to clear before the next release ships. If a problem happens in production and it does not turn into something the agent must pass, you are paying to discover the same problem twice.
Trust and governance are increasingly cited as major barriers to scaling agentic AI. Do you believe the technology is advancing faster than enterprises’ ability to supervise it, and what risks does that create?
I think that is exactly what is happening, and the gap is structural rather than a failure of effort. An idea can become a customer-facing agent in weeks. The operating discipline around that agent, the ownership, the evidence, the oversight, takes much longer, because it involves people and accountability and not only software.
The risk is that the gap stays invisible while it widens. An agent can give a customer a confidently wrong answer with no error, no failed transaction and no alert. Every dashboard looks green. Traditional operations depend on systems telling you when they are in trouble, and agents do not reliably do that.
I do not think the answer is to slow down. The companies that win here are going to move fast. The answer is to build the evidence and oversight that let you move fast with confidence. The more autonomy an agent receives, the more evidence you need that it can carry the responsibility.
When an autonomous agent makes a poor decision, who should ultimately be accountable: the developer, the business unit deploying it, the vendor providing the model, or the executive who approved its use?
Ultimately the company deploying the agent owns the outcome. Several parties are involved in building and operating the system, but the customer has no relationship with the model provider. The customer has a relationship with the company whose name is on the interaction.
That does not mean accountability sits with one person. It runs through the decision chain. The developer is responsible for how the system was built. The business decides what the agent is allowed to do. The vendor is responsible for the technology it provides. Leadership is responsible for making sure the company has the controls and the oversight to manage the risk at all.
The mistake is thinking that because the model made the decision, the model owns it. It does not. If an agent makes a commitment to a customer on your behalf, that commitment belongs to the brand. Customers understand this instinctively, and so do regulators.
AI agents can behave unpredictably when they encounter situations that were not anticipated during testing. How should enterprises test for these edge cases before agents are given access to customers, financial systems, or sensitive data?
You have to assume the agent will eventually encounter something it was not designed for. The question is what happens when it does.
So validate beyond the expected path. Give the agent ambiguous requests. Give it conflicting information. Give it incomplete context. Put it in situations where the right answer is to stop and escalate rather than keep going. Add the conditions of the real world, which in voice means accents, noise, poor connections and callers who change the subject halfway through. The objective is not to confirm that the agent works. It is to find out how it behaves when the conditions are not clean.
The more important point is that you have to validate the whole journey, not the agent in isolation. The model is usually not the problem. When something goes wrong, my first question is what context the model received. It may have been an outdated knowledge article, or two systems carrying conflicting policies, or a handoff that dropped what the customer had already explained. Every component can pass its own test and the customer journey can still fail in the seams between them.
That layer between the systems is the one we have spent years instrumenting, across 450 enterprises and more than 350 million customer journeys a year. Agentic or not, it breaks the same way. We also see agents built on more than 55 different vendors’ technology, plus every major contact center platform, which is how we know the pattern holds regardless of which model is underneath.
Before an agent gets access to something that matters, the enterprise should have evidence of what it does when things go right and when they do not.
How do you see AI testing evolving as companies move from deterministic software toward systems that reason, plan, communicate, and take actions across multiple applications?
Testing has to move from asking whether a system produced the expected answer to asking whether it achieved the right outcome.
That is a significant shift. An agent might take several different paths to solve the same customer problem, and those paths can change over time as the models and the knowledge behind them change. You cannot write a script for every possible interaction. You have to evaluate whether the agent understood the intent, made sound decisions along the way, and stayed within the boundaries it was given.
I want to be careful about one thing, because the industry is starting to get it wrong in an expensive way. Pre-launch testing matters more now, not less. It is what establishes whether an agent is ready. The argument that you can skip it and watch production instead is an argument for finding out in front of customers.
What changes is that pre-launch testing is no longer the end of the process. Production reveals conditions a controlled environment cannot reproduce completely, and what production reveals becomes a test the agent has to clear before the next release. Proof before launch, vigilance in production, and each one feeding the other. The agent running in month six should be measurably better than the one that launched.
Looking ahead, what will distinguish organizations that successfully build trusted AI workforces from those that remain stuck running small agentic AI pilots?
The organizations getting real return from agents are the ones that built an operating model around evidence. The ones that stall are usually not blocked by the technology. They are blocked because nobody can produce what the next level of approval requires. Legal asks a reasonable question, or the risk committee does, and there is no answer, so the pilot stays a pilot. The technology may be ready and the organization still cannot justify giving it more authority.
That is the difference between a pilot and an operating workforce. In a pilot, someone is always watching. In an operating model, each agent has a job you can state in a sentence. Its authority is limited and written down. Its performance is evaluated by something other than the team that built it. Production failures become release gates. More autonomy follows proof.
The second difference is ownership. In the companies that scale, the agent belongs to the business function it serves, with a named owner who answers for what it does. Where it remains an AI project owned by an AI team, it stays small, because no business leader will absorb the risk of something they do not control.
None of this is exotic. It is close to how a company already manages the people it trusts with real responsibility.
A pilot can run on an organization’s conviction. Scale requires evidence.
Thank you for the great interview, readers who wish to learn more should visit Cyara.












