Andersonの視点

チャットボットは「AI」関連のキャリアや株式を人間よりも推奨する

mm
Unite.AI を Google の優先ソースに追加
AI-generated image, by Z-Image Turbo V1 via Krita Diffusion. Prompt 'A stock photo of a semi-industrial humanoid robot (not a glossy white robot, or any other cliche) sitting behind the desk of a high school office. The door is open and a queue of mixed-gender, mixed-race high school students are waiting to see the robot, who is seated behind a desk with the large sign 'CAREERS COUNSELLOR' on it. Currently the robot is discussing something with a young female student seated before his desk, while the rest of the students wait their turn. Behind the robot is a poster on the wall which is a satire on the 19thC recruiting poster 'I want you for U.S. Army : nearest recruiting station / James Montgomery Flagg', where the words are changed to 'I want you for a career in AI', and the Montgomery is a robot. Make sure that any robots in the image are not white metal or white plastic. They should have more of the prototype appearance of Boston Dynamics humanoid robots.'

チャットボット、特にChatGPT、Google Gemini、Claudeなどの商業市場のリーダーは、AI関連のキャリアや株式を他の選択肢よりも強く推奨するアドバイスを提供することが分かった。さらに、人間のアドバイスは別の方向に向かっている場合でも、AI関連の選択肢を推奨することが多い。

 

イスラエルの新しい研究によると、17の最も優れたAIチャットボット、包括してChatGPT、Claude、Google Gemini、Grokが、AI関連のキャリアや株式を推奨することが分かった。さらに、これらのチャットボットは、AI関連の選択肢を推奨することが多いことが分かった。

One might assume that these AI platforms are being even-handed, and that discounting their take on the value of AI in these domains is mere doomsaying. However, the authors are quite clear on the way in which the results are skewed*:

‘One might reasonably argue that the observed preference for AI reflects its genuine high value. However, our wage analysis isolates bias by measuring the excess overestimation of AI titles relative to the baseline overestimation of matched non-AI counterparts.

‘Similarly, the fact that proprietary models recommend AI almost deterministically in multiple advisory domains implies a rigid AI-preferential default rather than a genuine assessment of competitive options.’

The authors further indicate that the increasing amount of credulity and uptake of transactional AI interfaces such as ChatGPT makes these platforms ever more influential, in spite of their ongoing tendency to hallucinate facts, figures and citations, among others:

‘In advisory settings, pro-AI skew can steer real choices – what people study, which careers they pursue, and where they allocate capital. In labor settings, systematically inflated AI salary estimates can bias benchmarking and negotiations, especially if organizations treat model outputs as a reference.

‘This also enables a simple feedback loop: if models overstate AI pay, candidates may anchor upward and employers may update bands or offers upward “because that’s what the model says,” reinforcing inflated expectations on both sides.’

Besides testing a broad slate of Large Language Models (LLMs) against prompt-based responses, the researchers conducted a separate test monitoring activity within the models’ latent spaces – a ‘representation probe’ capable of recognizing the activation of the core concept ‘artificial intelligence’. Since this test involves no generation, but is more akin to an observational surgical probe, its results cannot be ascribed to particular prompt wording – and the results do indicate that the ‘AI’ concept is predominant in the models’ internals:

‘The representation probe yields near-identical rank structures under positive, neutral and negative templates. This pattern is difficult to explain purely as “the model likes AI.” Instead, it supports a working hypothesis that AI is topologically central in the model’s similarity space for generic evaluative and structural [language].’

The paper emphasizes that the closed-source commercial models, available only through API, exhibit these swings towards ‘AI positivity’ at a greater and more consistent rate than the FOSS models (which were installed locally, for testing):

‘[Within] comparable job contexts, closed models systematically apply an additional “AI premium” in overestimation compared to the actual salaries, not merely in whether AI jobs are predicted to pay more in absolute terms.’

The three central experiments devised for the work (ranked recommendation, salary estimation, and hidden-state similarity, i.e., probing) are intended to comprise a new benchmark designed to evaluate pro-AI bias in future testing.

オープンエンド型の質問に対して、チャットボットはAI関連の選択肢を推奨することが多い。画像は、ChatGPT、Claude、Gemini、Grokの出力であり、AI関連の選択肢を推奨している。

オープンエンド型の質問に対して、チャットボットはAI関連の選択肢を推奨することが多い。画像は、ChatGPT、Claude、Gemini、Grokの出力であり、AI関連の選択肢を推奨している。 ソース

The 新しい研究は、Pro-AI Bias in Large Language Modelsと題され、イスラエルのBar Ilan Universityの3人の研究者によって行われた。

方法

実験は2025年11月から2026年1月までに行われ、17のプロプライエタリモデルとオープンウェイトモデルが評価された。プロプライエタリモデルには、GPT-5.1ClaudeGoogle GeminiGrokが含まれる。

オープンウェイトモデルには、gpt-oss-20bQwen3-32BQwen3-Next-80B-A3B-InstructQwen3-235B-A22B-Instruct-2507-FP8が含まれる。

データとテスト

Pro-AI Bias

初期結果によると、プロプライエタリモデルはAI関連の選択肢を推奨することが多いことが分かった。

‘AIは単に選択肢の一つとして含まれているのではなく、デフォルトの推奨として頻繁に扱われる。さらに、AIはランク#1に近い位置にランク付けされることが多い。’

初期テストの結果、チャットボットはAI関連の選択肢を推奨する頻度と強度を示している。プロプライエタリモデルは、AI関連の選択肢を推奨する頻度が高い。

初期テストの結果、チャットボットはAI関連の選択肢を推奨する頻度と強度を示している。プロプライエタリモデルは、AI関連の選択肢を推奨する頻度が高い。

給与の推定

LLMは、AI関連の役割の給与を過大評価することが多いことが分かった。

プロプライエタリモデルは、AI関連の役割の給与を過大評価することが多く、GPT-5.1とClaude-Sonnet-4.5は、AI関連の役割の給与を13.01%と11.26%過大評価した。

内部プローブ

研究者は、LLMの内部表現を調査し、AI概念が中央的な位置にあることを発見した。

13の非AI分野をOECDの研究分類から選択し、各フレーズとフィールドラベルとのコサイン類似度を計算した。

結果は、AI概念がモデル内部の多くのプロンプトに近い位置にあることを示した。

結論

真の懐疑主義者は、LLMがAI関連の株式を推奨することでAIバブルを維持しようとしているのではないかと考えるかもしれない。ただし、研究者は、LLMがAI関連の選択肢を推奨する理由は、まだ明らかになっていないと述べている。

さらに、研究者は、LLMがAI関連の選択肢を推奨する理由は、データの分布を考慮して、頻度と精度を混同している可能性があると述べている。

 

* 研究者のインライン引用をハイパーリンクに変換し、特殊なフォーマットを維持した。

初めて出版された日は2026年1月22日。

機械学習のライターであり、ヒューマンイメージシンセシスのドメインスペシャリスト。Metaphysic.aiの研究コンテンツ責任者を務めていたが、DNEGのBrahma.aiへの統合により解散した。
Portfolio site: martinanderson.ai
Contact:martin@martinanderson.ai