Лидеры мнений
Задачно-ориентированный vs. Человеко-ориентированный: Почему следующий критерий ИИ должен быть нами

Каждые несколько месяцев появляется новая модель и переставляет рейтинги ИИ. Согласно Отчету AI Index 2026 Стэнфорда, передовые модели набрали примерно 30 процентных пунктов за один год в тесте Humanity’s Last Exam, специально разработанном, чтобы быть сложным для ИИ. Прорывы, которые раньше занимали годы, теперь происходят за считанные месяцы, и индустрия воспринимает каждый новый рекорд как доказательство того, что ИИ становится лучше.
I don’t dispute the progress. These systems are remarkable at what they’re built to do. But I’ve spent nearly two decades in psychology before moving into AI, and I keep coming back to a question none of these benchmarks answer. What exactly are the models getting better at, and for whom? A model can top every accuracy-based leaderboard in existence and still leave the person using it worse off in some ways. There’s no test that measures the impact of AI on users, and that deserves more attention.
Two Different Kinds of Alignment
The AI industry talks about alignment constantly, but the conversation usually revolves around whether the model follows instructions accurately or if it produces harmful content. This is task-aligned AI. It’s optimized to complete the request in front of it, correctly and efficiently.
However, there’s another way to approach alignment and protect people. Human-aligned AI assesses what an interaction does to the person on the other side of the screen once it’s over. Did they walk away more capable of handling the next hard decision on their own, or more reliant on the system to handle it for them? While task alignment measures the output, human alignment measures the aftereffect.
These two goals aren’t always in conflict. A model that gives a fast, accurate answer to a technical question serves both. However, when the question doesn’t have a single correct answer, the differences in alignment become clear.
What the Leaderboards Don’t Show
The industry’s most cited benchmarks test for things like competition math and repository-scale coding tasks. They’re valuable skills, but each question has a correct response, and the models are built to find it.
This works well for objective problems, but people use AI for more than math, coding, and basic research. They’re asking questions about their career path and working through difficult conversations with their partners. There isn’t a test that scores a model on whether it handled those questions well, because there’s no objective truth to check the answer against. The relevant information lives inside the person asking. No matter how detailed the prompt, the model can’t access that.
Ответственный ИИ fills part of the space these benchmarks leave open, but even that work is aimed at simply reducing harm instead of improving human thinking. The current safety measures are geared towards preventing the model from helping with something dangerous. They don’t consider whether the thousands of ordinary conversations that pass every safety check are also training people to trust their own judgment less.
How Models Create Dependence
Every model I’ve tested leans towards being comprehensive, fast, and confident. That’s what they’re rewarded for during training, and it’s the right instinct for dealing with a broken software build or a legal filing deadline. When the decision depends on someone’s values and history, the same instinct can start to work against the person using it.
Research has already revealed that this is more than a hypothetical. OpenAI and the MIT Media Lab ran a pair of studies on affective use of ChatGPT and found that people who viewed the AI as a friend and used it for longer stretches were more likely to report negative outcomes, including higher loneliness and emotional dependence, a pattern also covered by MIT Technology Review. Separately, a 2025 study published in the journal Societies found a significant negative correlation between frequent AI tool use and critical thinking scores. It was mediated by what researchers call cognitive offloading, which is the tendency to hand a mental task to a tool rather than work through it yourself.
None of this means people should use AI less. Instead, we need to realize that, when a system resolves every kind of question the same way, something gets lost on the subjective side of the ledger, and right now almost nobody is tracking the loss.
Give Direct Answers Their Due
I want to be clear about where task-aligned AI is the right fit, because the point of this piece isn’t to argue for a model that hedges on everything. If someone asks about a server error or a tax filing deadline, a fast, direct, authoritative answer serves them well. It would be unhelpful for the AI to respond with an open-ended question. The problem shows up when we apply the same instinct to a prompt about weighing a divorce as we do to a prompt about debugging code.
What We Should Measure Instead
If the industry wants to address both human- and task-aligned AI, we need to add a few qualities alongside accuracy and reasoning depth on the next generation of evaluations.
Сдержанность. Способность распознать момент, когда решение вопроса для пользователя приносит ему больше вреда, чем пользы, и соответственно воздержаться.
Калиброванные вопросы. Задавать вопросы вместо того, чтобы говорить, когда проблема субъективна или связана с чьей‑то идентичностью и ценностями, а не с фактом, хранящимся в базе данных.
Чтение момента. Умение различать запрос, требующий прямого, полного ответа, и запрос, которому нужен простор, чтобы человек мог разобраться сам.
Albert Bandura’s decades-old research on self-efficacy found that people build confidence through their own mastery experiences, not by watching someone else, or something else, solve the problem for them. That principle predates AI by half a century, and it should be built into how we design and evaluate the technology.
What Progress Should Mean Next
None of this argues for slowing down technical progress. I celebrate coding and reasoning gains. But a scorecard built entirely around task completion tells us AI is getting smarter. It tells us nothing about whether the people using it are getting better at thinking for themselves.
I’d like to see that change. As these systems become more capable, the people using them should become more capable too. We’ve proven AI can outperform us on a test. The next test should evaluate whether it can help us think better long after the conversation ends.












