Thought leaders

Taak‑georiënteerd vs. mens‑georiënteerd: waarom de volgende benchmark voor AI ons moet zijn

mm
Voeg Unite.AI toe aan je voorkeursbronnen op Google

Elke paar maanden verschijnt er een nieuw model dat de AI‑leaderboards opschudt. Volgens het 2026 AI Index Report van Stanford hebben frontier‑modellen in één jaar ongeveer 30 procentpunten gewonnen op de Humanity’s Last Exam, een benchmark die speciaal is ontworpen om moeilijk te zijn voor AI. Sprongen die vroeger jaren duurden, gebeuren nu binnen enkele maanden, en de industrie beschouwt elke nieuwe topscore als bewijs dat AI beter wordt.

Ik betwist de vooruitgang niet. Deze systemen zijn opmerkelijk in wat ze zijn ontworpen om te doen. Maar ik heb bijna twee decennia psychologie gestudeerd voordat ik naar AI overstapte, en ik blijf terugkomen op een vraag die geen van deze benchmarks beantwoordt. Waar worden de modellen precies beter in, en voor wie? Een model kan elke nauwkeurigheid‑gebaseerde leaderboard overtreffen en toch de gebruiker op bepaalde manieren slechter af laten komen. Er bestaat geen test die de impact van AI op gebruikers meet, en dat verdient meer aandacht.

Two Different Kinds of Alignment

The AI industry talks about alignment constantly, but the conversation usually revolves around whether the model follows instructions accurately or if it produces harmful content. This is task-aligned AI. It’s optimized to complete the request in front of it, correctly and efficiently.

However, there’s another way to approach alignment and protect people. Human-aligned AI assesses what an interaction does to the person on the other side of the screen once it’s over. Did they walk away more capable of handling the next hard decision on their own, or more reliant on the system to handle it for them? While task alignment measures the output, human alignment measures the aftereffect.

These two goals aren’t always in conflict. A model that gives a fast, accurate answer to a technical question serves both. However, when the question doesn’t have a single correct answer, the differences in alignment become clear.

What the Leaderboards Don’t Show

The industry’s most cited benchmarks test for things like competition math and repository-scale coding tasks. They’re valuable skills, but each question has a correct response, and the models are built to find it.

This works well for objective problems, but people use AI for more than math, coding, and basic research. They’re asking questions about their career path and working through difficult conversations with their partners. There isn’t a test that scores a model on whether it handled those questions well, because there’s no objective truth to check the answer against. The relevant information lives inside the person asking. No matter how detailed the prompt, the model can’t access that.

Verantwoordelijke AI fills part of the space these benchmarks leave open, but even that work is aimed at simply reducing harm instead of improving human thinking. The current safety measures are geared towards preventing the model from helping with something dangerous. They don’t consider whether the thousands of ordinary conversations that pass every safety check are also training people to trust their own judgment less.

How Models Create Dependence

Every model I’ve tested leans towards being comprehensive, fast, and confident. That’s what they’re rewarded for during training, and it’s the right instinct for dealing with a broken software build or a legal filing deadline. When the decision depends on someone’s values and history, the same instinct can start to work against the person using it.

Research has already revealed that this is more than a hypothetical. OpenAI and the MIT Media Lab ran a een paar onderzoeken naar affectief gebruik van ChatGPT and found that people who viewed the AI as a friend and used it for longer stretches were more likely to report negative outcomes, including higher loneliness and emotional dependence, a pattern also covered by MIT Technology Review. Separately, a 2025‑studie gepubliceerd in het tijdschrift Societies found a significant negative correlation between frequent AI tool use and critical thinking scores. It was mediated by what researchers call cognitive offloading, which is the tendency to hand a mental task to a tool rather than work through it yourself.

None of this means people should use AI less. Instead, we need to realize that, when a system resolves every kind of question the same way, something gets lost on the subjective side of the ledger, and right now almost nobody is tracking the loss.

Give Direct Answers Their Due

I want to be clear about where task-aligned AI is the right fit, because the point of this piece isn’t to argue for a model that hedges on everything. If someone asks about a server error or a tax filing deadline, a fast, direct, authoritative answer serves them well. It would be unhelpful for the AI to respond with an open-ended question. The problem shows up when we apply the same instinct to a prompt about weighing a divorce as we do to a prompt about debugging code.

What We Should Measure Instead

If the industry wants to address both human- and task-aligned AI, we need to add a few qualities alongside accuracy and reasoning depth on the next generation of evaluations.

Terughoudendheid. Het vermogen om een moment te herkennen waarop het oplossen van de vraag voor de gebruiker meer schaadt dan helpt, en zich dienovereenkomstig in te houden.

Gekalibreerde vraagstelling. Vragen in plaats van vertellen, wanneer een probleem subjectief is of verbonden met iemands identiteit en waarden in plaats van een feit dat in een database staat.

Het moment lezen. Het onderscheid maken tussen een verzoek dat een direct, uitgebreid antwoord vereist en een verzoek dat ruimte nodig heeft voor de persoon om het zelf te verwerken.

Albert Bandura’s decades-old research on zelfeffectiviteit found that people build confidence through their own mastery experiences, not by watching someone else, or something else, solve the problem for them. That principle predates AI by half a century, and it should be built into how we design and evaluate the technology.

What Progress Should Mean Next

None of this argues for slowing down technical progress. I celebrate coding and reasoning gains. But a scorecard built entirely around task completion tells us AI is getting smarter. It tells us nothing about whether the people using it are getting better at thinking for themselves.

I’d like to see that change. As these systems become more capable, the people using them should become more capable too. We’ve proven AI can outperform us on a test. The next test should evaluate whether it can help us think better long after the conversation ends.

Chase Chick is CEO van Gray Sky AI, het bedrijf achter Vesela, en een erkende professionele counselor met meer dan een decennium ervaring in de gedragsgezondheidszorg. Hij heeft Pursuit of Happiness opgericht, een Texas counselingorganisatie die kinderen en gezinnen in het pleegzorgsysteem bedient. Zijn werk bij Gray Sky AI richt zich op mensgerichte AI en hoe technologie kritisch denken, menselijk oordeel en persoonlijke autonomie kan versterken. Chase brengt een praktijkgerichte blik op gesprekken over AI‑alignment, AI‑veiligheid, mens‑AI‑interactie en de impact van AI op hoe mensen denken en beslissingen nemen.