Interviews

Barb Hyman, Founder and CEO of Sapia.ai – Interview Series

mm
Add Unite.AI to your preferred sources on Google

Barb Hyman, Founder and CEO of Sapia.ai is an experienced business leader whose career has spanned law, management consulting, marketing, learning and development, human resources, and technology. She began her career as a solicitor at Herbert Smith Freehills before moving into consulting with Boston Consulting Group, where she progressed to Project Leader and later returned in senior roles overseeing learning and development, HR, and marketing across Australasia. Hyman also served as Head of Marketing and Sponsorship at the Museum of Contemporary Art in Sydney and as Executive General Manager of People & Culture at REA Group. She founded Sapia.ai in 2018, drawing on her experience in organizational leadership and talent management to rethink how companies assess and hire people using artificial intelligence.

Sapia.ai is an AI-powered hiring technology company focused on structured, conversational candidate assessment. Its platform uses chat-based interviews to evaluate candidates consistently at scale, helping employers assess skills and suitability while reducing reliance on traditional CV screening and other signals that can introduce bias. Sapia.ai says its technology has assessed more than 10 million candidates and combines structured assessment science with AI-powered talent intelligence, personalized candidate feedback, and integrations with existing applicant tracking and HR systems. The company also emphasizes explainability, fairness, and enterprise compliance, with its broader platform spanning AI interviews, job analysis, interview scheduling, talent intelligence, and AI-powered career coaching.

Before founding Sapia.ai, you built a career spanning consulting at Boston Consulting Group, law, marketing, learning and development, and eventually People & Culture leadership at REA Group. What experiences from that journey convinced you that hiring needed to be fundamentally rethought, and what problem did you originally set out to solve with Sapia.ai

The current way we assess people is broken — especially in hiring.

We rely on CVs. We don’t see the whole person. We make our most consequential decisions — who to hire, who to promote — on gut feel and proxies. Mirror hiring is rampant: since when did the school or university someone attended become a proxy for intelligence? It’s not. It’s a proxy for advantage.

Some of this relates to my personal story. I’m an immigrant and I know I landed in a fortunate category of immigrant: I’m white and I have a name people find easy to say. I come from a family where nobody went to university; my dad was a shopkeeper. And yet I’ve had extraordinary opportunities and five careers later, here I am. When I think about the two billion workers in the world, I think about how many of them have exactly the talent, the drive, the critical thinking I’ve been credited with, but are never seen because we’re blinded by the prism of our own experience. Bias is a wonderful shortcut for making a fast decision, and it’s exactly why hiring keeps rewarding the already-advantaged. It never felt fair or rational.

The other half of the story came from consulting. After my MBA, I moved into a world where you never walked into a boardroom and made a recommendation without real research behind it. And yet the highest-stakes decision any company makes (who joins, who leads, who shapes the culture) is still made on gut, or on a proxy as thin as a university name. People aren’t on the balance sheet, but in professional services especially, they are the asset. That disconnect seemed absurd to me.

The CV itself hasn’t fundamentally changed since Leonardo da Vinci invented the format. Wikipedia lists more than 180 documented human biases. A third of North American candidates who finish a job application are left without an answer two months later. None of that is a talent problem: talent is equally distributed but opportunity is not.

What I set out to fix with Sapia wasn’t just a decision — it was from the experience of looking for a job. My team laughs at me for this, but I’ve always thought of hiring like dating. It’s a relationship. You’ll spend more time with the people at your job than with your partner. So why would you make that decision on anything less than the same rigor — the data — we insist on for every other decision that matters?

You’ve argued that most organizations approach AI adoption as an efficiency project when it should begin as a trust project. What does “trust” actually look like in practice, and how should companies measure it alongside more traditional metrics such as productivity, cost savings, and adoption?

I learned this as a CHRO: if you’re not transparent about why you’re making a change, trust is the first casualty. And I think most companies get AI adoption backwards. They measure the productivity gain, e.g. the hours saved, the cost per hire, but they never stop to measure the thing that actually determines whether any of it sticks: do your people trust how you’re using AI, and why?

Trust is the substrate of your culture. It’s not a soft metric sitting next to the real ones — it’s the thing the real ones are built on. Think about it the way you’d think about any relationship. High trust gives you degrees of freedom. Low trust makes everything smaller, less buoyant, less relaxed. And nobody does their best work when they don’t feel safe. Organizations are no different.

Yet so much of what’s been done with AI in the workplace has actually weaponized people’s own data against them, for instance scraping someone’s social media to infer whether they’re a good hire, a good risk, a bad risk. I genuinely cannot imagine the meeting where that decision got made, or how anyone thought it wouldn’t corrode culture the moment people found out.

So if you’re serious about AI adoption, trust has to be a metric from day one, sitting right alongside productivity and cost savings. Ask people directly: do you understand why we’re using this? Do you believe it’s being used fairly? If you can’t answer yes to both, the efficiency gains you’re celebrating are borrowed against a culture you’re quietly spending down.

One of the biggest concerns around AI in hiring is that models can learn from historical employment data and potentially reinforce the same human biases organizations are trying to eliminate. How can companies prevent AI from simply automating existing hiring biases, and what safeguards should be built into these systems from the outset?

This is exactly the trap most AI-in-hiring conversations miss and the reason I built Sapia the way I did. If you train a model on historical hiring data (who got hired, who got promoted, who got that first interview) you’re not removing bias, you’re industrializing it. You’re teaching a machine to replicate every mirror-hiring decision, every “culture fit” call, every gut-based shortcut that got you the workforce you already have. That’s not innovation. That’s bias at scale, just faster and with a veneer of objectivity that makes it harder to challenge.

So the safeguard has to start further back than the algorithm. We don’t train on CVs, photos, names, universities, or any of the traditional proxies that correlate with privilege rather than potential. If the input carries the bias, no amount of clever modeling downstream fixes it. It’s the “garbage in, garbage out” problem, except the garbage is decades of human prejudice dressed up as a resume.

The second piece is what you actually measure the model against. It’s not enough to say “the AI works,” you have to constantly test for adverse impact across gender, ethnicity, age, disability, every protected characteristic you can measure continuously, not as a one-time audit before launch. Bias isn’t a bug you fix once. It’s something you have to keep watching for, the same way you’d monitor any system that makes consequential decisions about people’s lives.

And third — this is the one people skip — you need outside eyes on it. Independent, third-party audits of your models, published methodology, real transparency about what the system is and isn’t doing. If you’re only willing to have your AI hiring tool checked by the people who built it and are financially motivated to defend it, you haven’t actually solved the trust problem, you’ve just relocated it.

The “algorithm did it” is not an excuse, but an admission you didn’t build the safeguards in the first place. Remember, AI doesn’t remove human responsibility for fairness.

Hiring is an unusually high-stakes application of AI because an algorithm can directly influence someone’s career. Where should organizations draw the line between decisions that can be automated and those that should always remain under meaningful human oversight?

The line I draw is this: AI can absolutely help you see a person more clearly, but it should never be the one who decides their fate alone. There’s a real difference between augmenting a decision and outsourcing it. Hiring is one of the few places where getting that distinction wrong actually changes someone’s life.

So where does automation earn its place? In surfacing signals that humans are bad at seeing consistently and fairly, for instance reading thousands of interview responses without fatigue, without having a bad day, without unconsciously favoring the candidate who reminds them of themselves at twenty-five. That’s a job AI can do better than a tired hiring manager on their eighth interview of the day. Screening for genuine soft skills, structuring insight so every candidate is assessed on the same criteria, that’s where the machine adds real value because we’re not built to be that consistent.

The final call — who gets hired, who gets rejected, who gets a second look — has to sit with a human who can be held accountable for it. Not as a rubber stamp on whatever the model outputs, but real, meaningful oversight: someone who understands what the system is measuring, can question a result that looks wrong, and can override it. If a human can’t meaningfully challenge the AI’s recommendation, it’s not an oversight but a human alibi.

I’d also draw the line at anything irreversible or opaque. If a candidate can’t get an understandable explanation for why they didn’t progress, the system has failed regardless of how “accurate” it is. And if the model is making decisions nobody in the company can actually explain — including the people who built it — that’s not a hiring tool, that’s a black box with someone’s career inside it. The test I’d give any organization is simple: could you sit across from a candidate you rejected and explain, in plain language, why? If the honest answer is “I don’t actually know, the model decided,” you’ve automated the wrong part of the process.

Sapia.ai distinguishes between generative AI used for open-ended conversations and more deterministic systems used for scoring candidates. Why is that distinction important, particularly when AI outputs need to be consistent, auditable, and defensible?

People hear “AI” and assume it’s one thing, but in hiring that conflation is exactly where things go wrong. We deliberately split the job in two, because open-ended conversation and high-stakes scoring have completely different requirements, and pretending they’re the same technology is how companies end up with systems they can’t explain, defend, or trust.

Generative AI is brilliant at the conversational layer — having a natural, human-feeling exchange with a candidate, asking good follow-up questions, making the experience feel like a conversation rather than a form. That’s where you want flexibility, nuance, even a bit of unpredictability, because real conversations aren’t scripted.

But the moment you move from talking to a candidate to scoring them, everything about what you need from the system flips. Scoring has to be consistent: the same answer should get the same assessment whether it’s candidate one or candidate one thousand, whether it’s a Monday morning or a Friday afternoon. It has to be auditable, which means you need to be able to go back and show exactly why a candidate got the score they got, not just trust that the model “figured it out.” And it has to be defensible because someone’s livelihood is on the other end of that score, and “the generative model felt like this was a good answer” is not a sentence that survives contact with a regulator, a journalist, or a candidate who wants to know why they didn’t get the job.

Generative models, by their nature, are probabilistic — ask the same question twice and you can get subtly different outputs. That’s a feature for conversation and a serious liability for scoring. At Sapia, we use deterministic systems for the part of the process that actually decides someone’s outcome. We model where the same input reliably produces the same output, where we can trace the reasoning, and where we can prove that two similar candidates are treated equivalently.

It’s really a separation of concerns: let generative AI be great at being human-like, and let deterministic, auditable systems carry the weight of a decision that has to be fair, explainable, and defensible under scrutiny.

Explainability is often discussed as a requirement for responsible AI, but many advanced models remain difficult for users to understand. How much explanation should a candidate or employee reasonably expect when AI has influenced a hiring decision?

I think candidates deserve more than most companies currently give them, and honestly, more than most companies think they’re capable of giving. “The algorithm decided” is an admission of failure dressed up as sophistication.

Here’s where I land: a candidate should always be able to understand, in plain language, what they were assessed on and roughly why they landed where they did. Not the underlying math, not a printout of model weights. What they want and what they’re owed, is something closer to what a good manager would tell them face to face: these are the qualities we were looking for, this is where you were strong, this is where you weren’t the closest match this time. That’s a completely reasonable bar, and any system that can’t clear it shouldn’t be making decisions about someone’s livelihood in the first place.

So for us, explainability has to be baked into the architecture from day one, because you cannot explain what you cannot trace.

Where I’d push back a little is on the idea that explainability means candidates need a technical education in how the model works. That’s a false standard, and it’s often used as an excuse when really it just hasn’t been designed to be explained. A good doctor doesn’t walk you through the biochemistry of a diagnosis; they tell you what they found and what it means for you. AI in hiring should hold itself to the same standard of human-readable honesty.

The real test is the same one I’d apply to human oversight: could a candidate ask “why didn’t I progress?” and get an answer that’s specific, honest, and actually theirs. If your AI can’t produce that answer, you haven’t built an explainable system. You’ve built one you’re hoping nobody asks about.

Companies frequently focus on whether employees and managers are using an AI tool, rather than whether they actually trust its outputs. What warning signs suggest an organization has achieved AI adoption without achieving genuine confidence in the technology?

Adoption is the easiest metric to fake, and that’s exactly why so many companies mistake it for success. If I tell you everyone in the business is “using” the AI tool, that tells you almost nothing about whether they believe in it — it might just mean it’s mandatory, or it’s the path of least resistance, or nobody’s found a way around it yet.

The first warning sign is workaround behavior. If people are using the AI tool but also quietly doing the old process in parallel — double-checking its output manually, keeping their own separate notes, running a shadow system just in case — that’s compliance with an insurance policy attached. Usage numbers look great. The truth is everyone thinks the tool might be wrong.

The second is silence. In a genuinely trusting culture, people challenge the system when something looks off: they ask questions, push back, flag results that don’t feel right. When adoption is real but trust isn’t, you get the opposite. People stop questioning it but because they’ve decided it’s not worth the fight. That silence gets misread by leadership as satisfaction. It’s actually a resignation.

The third warning is a gap between what people say in a survey and what they say to each other. I’d be far more worried about a company where the official adoption dashboard is green across the board but the real conversation in the break room is “yeah, we just do what it says because that’s easier than arguing” — that’s a culture that has not converted.

And the deepest warning sign and the one that matters most to me: when people stop bringing their own judgement to the table at all. That’s abdication. Real trust in a tool means people still feel empowered to disagree with it when their own expertise tells them something different — and feel safe doing so. If your managers have gone quiet and just defer to whatever the AI says, you’ve built dependency.

So the real test isn’t “are people using it” but “would people choose to keep using it if it were optional, and would they tell you honestly when it got something wrong.” Most companies have never actually asked either question. That’s usually the tell.

AI hiring tools can potentially evaluate candidates at a scale and consistency that human recruiters cannot, but hiring also involves context and nuance. Where do you believe AI can outperform human judgment today, and where are humans still essential?

AI can be a great tool at assessing a large pool of candidates regardless of their backgrounds and experiences, all while providing tailored assessment for candidates with specific needs. A human recruiter might be more inclined to favor candidates from a top-ranked university or without previous health conditions. AI conducts a fair, equal assessment of all candidates in a way that humans can’t do when they face hundreds of applications. That said, human judgment is still essential to evaluate whether a candidate will fit into a company culture or meet the standards required by clients.

As AI becomes more deeply embedded in recruitment, coaching, performance management, and other workforce decisions, how should accountability be structured when an AI-assisted decision produces an unfair or incorrect outcome?

Accountability in AI remains when humans are the final decision-makers while the recruitment workflow includes transparency, explainability, and rigorous bias testing. We launched the FAIR™ model to help drive responsible AI practices. It recommends:

  • Unbiased: Test for bias and use the right statistical measures — such as the 4/5ths rule, error-rate parity, and other relevant bias tests — to check that outcomes are fair.
  • Validity: AI-driven evidence is based on actual data and linked to the outcome it is intended to predict.
  • Explainability: Have clear documentation and tools to interpret results, what the AI has identified about a candidate and how that has contributed to their score.
  • Inclusivity: This means making sure candidates are treated fairly throughout the entire process, from application through to decision-making. It’s an end-to-end view of the hiring experience — including the people, processes, technology, and AI working together.

Looking ahead, do you expect AI to primarily become a decision-making layer inside HR systems, or will its more important role be conversational, helping candidates and employees better understand opportunities, develop skills, and navigate their careers?

Even the most advanced technologies require human touch, human judgment, and human oversight. AI is there to support the process, making sure every candidate has a fair, equal chance to be considered for a position. At Sapia, we like to call this a “human-in-the-loop approach” where we’re combining human and machine intelligence to create smart chat experiences that are predictive, responsible and personalized. To do so, we bring our customers onboard, so we constantly validate for accuracy and fairness, all while teaching our AI models to evolve in time.

Our human approach also translates in what we offer to candidates. We recently launched Phai, a career coach that acts as a mentor, providing jobseekers with personalized guidance, strengths mapping, and interview preparation to help them confidently face their next opportunity. It’s a positive tool to keep a candidate motivated beyond the interview outcome.

Thank you for the great interview, readers who wish to learn more should visit Sapia.ai.

Antoine is a visionary leader and founding partner of Unite.AI, driven by an unwavering passion for shaping and promoting the future of AI and robotics. A serial entrepreneur, he believes that AI will be as disruptive to society as electricity, and is often caught raving about the potential of disruptive technologies and AGI.

As a futurist, he is dedicated to exploring how these innovations will shape our world. In addition, he is the founder of Securities.io, a platform focused on investing in cutting-edge technologies that are redefining the future and reshaping entire sectors.