Interviews
Yuri Gubin, CTO at DataArt – Interview Series

Yuri Gubin, CTO at DataArt is a veteran technology executive and software architect who has spent more than 18 years with DataArt, progressing through roles spanning software architecture, solutions architecture, cloud technology, innovation, and executive leadership before becoming Chief Technology Officer in March 2026. His work has focused on solving complex technology challenges across industries including financial services, healthcare, travel, and IoT, with particular expertise in cloud computing, AI, data platforms, and enterprise software architecture. Before becoming CTO, Gubin served for more than five years as DataArt’s Chief Innovation Officer and has been a member of the company’s Board of Partners since 2021. He is also a professional member of the Forbes Technology Council, participating in its AI and Cloud Computing expert groups, and serves as a Technology Advisor to Girls Who Code, where he advises on architecture, data protection, platform governance, and technology policy. DataArt currently lists him as its Chief Technology Officer based in New York.
DataArt is a global software engineering and data and AI transformation company founded in New York in 1997. The company has grown to more than 6,000 technology professionals operating across 20+ countries and works with more than 400 clients, providing services in areas including artificial intelligence and machine learning, data and analytics, cloud transformation, custom software engineering, cybersecurity, and legacy modernization. DataArt works across sectors such as financial services, healthcare and life sciences, travel, media and entertainment, and retail, and maintains technology partnerships with platforms including AWS, Google Cloud, Microsoft Azure, Snowflake, and Databricks. In 2025, the company announced a $100 million, three-year investment in its data and AI capabilities, followed in 2026 by the launch of Artisyn, an AI-enabled operating model designed to incorporate AI agents, reusable accelerators, governance, security, and compliance into enterprise software development.
You’ve spent nearly two decades at DataArt, progressing from software architect and solutions architect to Chief Innovation Officer and now CTO. How has that journey shaped the way you distinguish genuinely transformative technologies from cycles of hype, and how does it inform your “skeptical optimism” toward AI today?
We have seen many different waves over the years, including the rise of cloud and mobile, different generations of AI, automation, DevOps and SRE, and I’ve been coding, architecting and advising our clients on many of these topics throughout that time. What I realized is that, yes, you can do almost anything with technology, and technology is quite powerful, but the devil is in the detailsand you need to know what you’re doing for it to make sense and work.
I have seen cloud environments become more and more expensive, AI models that don’t perform the way you think they will, and poorly implemented attempts to automate release cycles. I’ve seen the impact of both good decisions and bad decisions, so whenever something new comes along and you read all the announcements, promises and hype, I go back to the same premise: almost anything is possible with technology, but you need to know what you’re doing.
You get a good grasp of a technology through R&D and, importantly, through real-life projects, because that is how you learn what is possible, what isn’t, and where things can go wrong. You take those lessons from every engagement, talk to your peers, other architects and analysts, and try to understand whether there are patterns and whether you can create some kind of system around them. Eventually, that becomes guidance, and then you see whether the decisions you thought were good decisions really produce good outcomes.
That is where the skeptical optimism comes from. Whatever technology promises, you still need to know what you’re doing, and that knowledge comes from experience, collaboration, and a continuous effort to learn, get better and create some kind of system behind the hype.
Enterprise AI seems to be moving from a phase of encouraging experimentation to deciding which experiments actually deserve to scale. What signals tell you that an AI use case is ready for broader deployment, and what are the warning signs that a company is scaling too early?
I use two methods to understand whether we can scale something or whether we need to do something else: the adoption curve and the learning curve.
To understand if an AI use case is working, you need to give it some time and understand what value it brings and what the user journey looks like, because then you can see the ups and downs rather than just the immediate ‘wow’ effect in a particular team or workflow. You need to see what happens to the same people a few weeks later. Are they still using it? Are they still happy with that use case, that automation or that AI skill they created, or was it just a blip that really shouldn’t be scaled?
Some of these things can only be validated over time. There will always be the first pioneers, usually the most technically savvy people and the ones who are very curious, and then you need to try it with other segments, with those who follow the early adopters and then the early majority. Once it proves itself there, yes, you can start scaling it and expanding that use case to other departments.
Every major model release can create pressure inside an organization to immediately give employees access to the latest capabilities. How should technology leaders evaluate whether a new model represents a meaningful improvement rather than simply generating another wave of experimentation and cost?
Here’s my skeptical optimism again. Assume that you already have a model in place and several thousand people using AI daily, with different models and tools already available. When a new model comes out, because of the hype and natural curiosity, you can expect that everyone will want to experiment with it, which is good, but that experimentation might not necessarily be guided or oriented toward any specific results, and sometimes you won’t even be able to measure the difference.
At scale, that matters. It’s not just one or two people poking around to see how the new model performs against the old one. It can be thousands of people spending time experimenting when, for a particular use case, the outcome may not be that significant. At the same time, if something does work really well, the learning about what works within your organization may not be clearly explained or visible to everybody.
That’s why the first group evaluating a new model should not be everyone in the organization. It should be an R&D group working closely with the relevant teams, as well as legal and security. We evaluate the model comprehensively, do a quick assessment, and then bring it to the broader audience with some comments and guidance around security, compliance and technology. With new models and major updates arriving constantly, you need to have this model and mindset in place. It really isn’t a one-time or one-off exercise.
DataArt has created a cross-functional “AI SWAT” involving technology, legal, compliance, InfoSec, and other teams. How does this group operate in practice, and what types of risks or questions need to be resolved before a new AI tool is approved for broader use?
Since its inception, I think we have set different objectives for this group roughly every four or five months. We change the priority, the goal and sometimes the mission, and a lot of these objectives are around AI. It can be upskilling the workforce, go-to-market and new capabilities, partnerships, or enabling AI more broadly within the organization and across the ADLC.
The exact topics evolve over time, and I think that’s healthy because you constantly need to revisit your own strategy, validate your assumptions, and understand whether you need to pivot and what the next theme for the team should be.
The group involves representatives from different departments, and one of its purposes is simply to keep everyone informed. Whenever there is a new announcement, question or opportunity, someone can bring that topic to one of our regular meetings. Even if it looks like a technology question relevant only to a narrow team, nowadays these topics can have implications for many parts of the organization.
That’s why, when we’re evaluating a new partnership, tool or accelerator, we discuss it openly so everyone understands where things are going and has a chance to ask questions or provide oversight. For a new AI tool, technology cannot evaluate it in isolation. Security, legal and compliance also need to understand how it handles company or client data, what constraints apply, and whether it can be used safely at scale.
Sometimes the AI SWAT team also works on specific programs, such as upskilling, where we set targets, map out roadmaps and decide how different groups will be onboarded. That is really how it works: keeping people informed, working together on specific programs, and providing visibility to the board into what is happening with AI across the company.
You’re seeing very different attitudes toward AI-assisted software development, with some organizations actively scaling agentic development while others still prohibit AI-generated code. What explains this divide, and what needs to change before more risk-conscious enterprises become comfortable with AI playing a larger role in software engineering?
Probably what drives the difference between those who say no and those who say yes is their risk appetite and their attitude toward ambiguity and uncertainty. What helps both types of organizations is continuous education, experimentation and evaluation. Even among many organizations we work with that embrace AI and incorporate it everywhere, there are still challenges around measuring the outcome and impact. To be honest, the question of how you measure the impact of AI and how you evaluate the performance of a team sometimes comes almost out of nowhere, as if nobody had really thought about it before.
Once you start evaluating an AI initiative more comprehensively, you begin to understand the impact and the value it is actually giving you, which leads to better decisions about where the technology makes sense. For companies that say no to AI, there still needs to be a continuous process for reviewing what the technology can do and where it stands now. You don’t want a decision made three years ago to remain company policy simply because nobody revisited the assumptions behind it.
Agentic AI makes it increasingly easy for individual teams to create their own agents, potentially resulting in multiple agents performing nearly identical tasks. At what point does experimentation become agent sprawl, and what kind of governance layer is needed to manage ownership, permissions, duplication, and lifecycle management?
When we see a typical scenario where an AI license is given to every developer and experimentation becomes unguided, everyone starts creating their own stuff and working in their own way. Typically, that leads to underperforming teams, missed expectations, lagging quality and rising expenses. The bottom line is that it doesn’t do what everyone expects it to do, the quality is bad, and it becomes expensive. To mitigate that, it needs to be a team effort that is part of a broader department or organizational effort, and that is where governance comes in.
At the project level, you can agree on the knowledge base and context, as well as the use cases where you start using AI. Then you create skills and agents that are part of the development workflow and that everyone can reuse, so you accumulate knowledge and best practices rather than recreating them every time. That project-level effort should then be orchestrated by something like an enterprise architecture board, a technology group, the CTO, or a team responsible for AI adoption. You want to reuse agents that work well, make sure the process is solid, and have that work across the organization rather than turning into chaos and noise.
So I think it has to be a synchronized effort at the project level, maybe the program level, and then at the department and organizational levels as well.
Token consumption and inference costs can appear relatively small during a pilot but become significant when AI systems are deployed across thousands of employees or autonomous agents. How should enterprises think about AI cost management, and do you expect something resembling FinOps to emerge specifically for AI workloads?
I will start by saying that an almost ideal scenario is when AI costs go up, reach a plateau, and then start decreasing slightly over time. That tells you that you can forecast, control the costs, understand what you are actually spending on AI, and see the outcomes of the decisions you make. The bad situations are when costs keep going up and down because that usually means something is not sustainable, or when costs rise and then drop completely because adoption may not be happening, something does not work, or people are using something else and you simply do not see it.
So FinOps is a thing, and AI FinOps is a thing as well. Some techniques are very technical, while others are quite simple. It can be as basic as choosing the preferred model so you do not always rely on the most expensive one, and piece by piece those decisions start saving money. At the same time, knowing how to save and control costs is only half of the equation. FinOps, as I see it, is a discipline and methodology that also involves product and business leaders because you need to define what you are measuring when evaluating AI efforts.
So yes, I think AI FinOps is a good topic for the equivalent of an AI SWAT team to discuss: how much you spend, how much you get back, how you control it, and where the opportunities are.
Many companies are being asked to demonstrate ROI from AI even though they never established a reliable baseline for how productive their teams were before AI was introduced. What should organizations actually measure if they want to determine whether AI is creating meaningful business value?
Regardless of your attitude toward AI or where you are right now, maybe you are already using agents left and right or maybe you are thinking that next year you will start using AI, establishing a baseline is absolutely a must-have nowadays.
There are several classes of metrics. Some are subjective, and that can simply be feedback from your developers or employees because you are working with people and it is important to understand how they perceive the value of AI. More objective measures can start with mechanical or synthetic metrics, although I would urge everyone not to bind themselves too closely to them. I mean things like code commits or story points. These metrics show that work was happening, but they do not really show the value or impact.
What makes more of a difference are metrics that explain how quickly or how well the work was delivered. Think about DORA metrics such as lead time or MTTR, how quickly you can recover from a failure, how quickly you can fix a bug in production, or how those measures change over time. A number at one point does not tell you the trajectory. One of our architects mentioned recently that, in software development, a good metric might also be how reliable estimates are as AI adoption grows, because that says something about the sustainability of these efforts and how productive teams really are. You also need to keep track of costs because if you only talk about benefits without understanding what it costs to achieve them, you do not have the full picture.
Outside software development, I think about it in a similar way. In every workflow or process, there is some unit of work and some definition of done. Whether you are processing claims, reviewing paperwork or handling customer requests, define what you are delivering and then measure how long it took before AI, how quickly and how well you can do it now, and what it costs. That gives you a good starting point for both the baseline and the framework of metrics.
DataArt has been embedding AI throughout the software delivery lifecycle through initiatives such as Artisyn. As AI takes over more implementation, testing, and workflow tasks, which parts of software engineering become more valuable for humans, and which skills risk becoming less important?
You can only use AI effectively in development if you still remember what the definition of good is. You need that expertise to guide your agents, review the outcome, set constraints and define the rules. You need to understand what best practice is and what good architecture should look like because, without that, you may not know what is being developed, and the value of this kind of expertise is rising very, very significantly.
Understanding architectural patterns matters, as does understanding what is appropriate in a certain industry, application or class of solution. You need to know what kind of architecture is good right now and what kind will still be good when the solution scales, because sometimes the same architecture does not work throughout the lifetime of a solution or platform.
That balance of what is appropriate for a particular solution is the human part. This is the taste, the artisanship behind services and software development. You have to know what you are doing, and that also comes from understanding the client and the industry.
Which skills are less important? It’s really hard for me to say that, although maybe how quickly you can type code. I’m joking, but code can now be created much, much more quickly, and specific knowledge of a particular library or language can also be learned much faster with AI.
I’ve seen .NET developers retrained as Java developers very quickly, and five or ten years ago I would have said that doing that at scale was almost impossible. Nowadays, you can. A strong senior developer can increasingly move between languages because what really matters is their understanding of technology, architecture, solution best practices, SDLC and ADLC.
As enterprises move from dozens of AI pilots toward production systems that can independently take actions, where should accountability ultimately sit when an AI agent makes a costly mistake: with the developer, the business owner, the model provider, the governance team, or some combination of them?
I like the idea of blameless collaboration and shared responsibility because everyone in the organization contributes to best practices, architectural frameworks and solutions. Even if one developer creates code with AI or without AI, another developer reviews it, team leads provide guidance, architects provide the architecture and constraints, and the governance team contributes to decisions about budgets, timing and release. Everyone is involved in some way.
Very often, when something goes wrong, it is the process that does not work, so in that sense accountability is shared across different roles. But if you simply say accountability is shared and therefore blameless, that is not enough. It still needs to be broken down into specific responsibilities.
Developers are accountable for the code they submit as a pull request, and they need to understand what is going on there. Architects are accountable for the decisions they make and for the architectural decisions provided to agents and developers. The platform team is accountable for the reliability of the solution regardless of who or what created a particular line of code.
So the accountability is there, but you need to define it granularly by team, role and department. What you cannot do is stop the analysis at “AI did this.” You need to ask what controls, tests or oversight allowed that failure to reach production.
If a lack of unit tests allowed bad code to be pushed into production, or a lack of oversight and review allowed it to happen, you cannot defer that accountability to AI. You also cannot simply blame the model provider or cloud provider for every bug or outage.
Thank you for the great interview, readers who wish to learn more should visit DataArt.












