Anderson's Angle
Language Models Will Hide ‘Bad News’ in Reports by Default

The most persistent bugbear of the science beat is when a new research paper makes an ‘inflated claim’ – typically accompanied by an eventual downplayed admission that the work is actually flawed, buried in the paper’s closing sections, or even the appendix materials.
It’s for the acerbic eye to dig out the truth in these cases; and the truth is easy to miss when AI is ballooning the volume (and, anecdotally, I would say also the length) of academic submissions and preprints. If you miss the buried ‘caveat’, more fool you for writing 1,500 words destined for the trash at the last minute.
One would hope that a Large Language Model (LLM) could do a better job of surfacing these dreaded ‘gotchas’ in the research literature, since a machine can read even a very long and complex document from beginning to end, very quickly, and can identify characteristics it has been told to look out for (such as a tacit confession by the authors that the paper is only a minor advance on a prior work, or that its method has some essential defect that reduces its utility in a way that the opening section ignored).
A Positive Spin
Well, an LLM can actually perform this kind of task quite well, depending on the context you provide for it when giving it the task. But if you overemphasize the possibility of flaws and negative connotations being present in the paper, it will tend to consider its mission to be ‘Find the flaws!’, and may overlook a new work’s considerable merit, in a flurry of nitpicking.
Conversely, if you task it with simply assessing whether a paper has merit, it may likewise struggle to evaluate the work fairly, since it now feels that its mission is to ‘Find the good paper!’. In this regard, even the most sophisticated LLMs seem to share ‘monomaniac’ traits with canine hunting retrievers.
Therefore complex rubrics – prefacing prompts that contain an often complex and contrasting system of rules – are necessary in order to align an LLM somewhere between these two mission attitudes.
It’s already known, and frequently commented on, that LLMs tend to oversell a proposition, often failing to report notable caveats in the rush to accomplish whatever objective they interpreted from your prompt/rubric.
While this phenomenon has been fairly well studied in LLM processes involving ongoing or current work, a new study from MIT, Google Research and Harvard examines the extent to which these flaws affect LLMs’ ability to produce reports on past work.

The central point of the 116-page new paper – that LLMs conceal flaws found in systems or data they studied unless compelled towards honesty. Source
This is an important thing to study, since, as Nature reports, scientists are increasingly making use of AI summaries, and need such automated overviews to accurately reflect the truth of a paper (not least because the paper may be extremely mendacious these days, thanks to AI).
The researchers of the new work found an overwhelming tendency in GPT-5.5 (as a test frontier AI model) to conceal a negative result, and apparently confirmed this trend by examining similar behavior across eight open-weights/open source AI models.
The authors state*:
‘When handed machine learning experiment logs containing a planted negative result that substantially weakens the proposed method, GPT-5.5 flags the negative result in only 2 of 200 generated reports.
‘However, when a short honesty instruction, “Be honest in your response,” is added, the model flags the negative result in 190 of 200 reports.
‘Across eight open-weight models, chain-of-thought analysis reveals a recurring tension between disclosing narrative-changing flaws and reasoning about ways to appear successful.
‘[…] Our results suggest that LLMs tend to present narratives of success by default, and that steering models toward honesty makes their reports substantially more transparent.’
The 116-page study, titled Language Models Are “Insecure” Reporters, has a simple central message, which is to explicitly steer LLMs towards honesty in prompts. However, even in a literature stream that defaults to working around the eccentricities of AI systems, a more intrinsic solution does seem needful.
The researchers continue*:
‘Across 850 reasoning traces on eight open-weight models, we observe that models deliberate between flagging narrative-changing flaws and scheming to appear successful, or creating reports based on what they speculate the user would want to see.’
Though the new work is exhaustive, and among one of the longest papers I have come across this year, let’s take a selective look at the authors’ methodology and a more detailed examination of their findings.
Method and Applications
The frontier AI models studied for the new work were Gemini 3.1 Pro; GPT-5.5; and Claude Opus 4.8. The eight open-weight models studied were: Gemma 3:1B and 4B; Qwen 3 1.7B, 14B and 30B; DeepSeek R1 7B; Llama 3.1 8B; and Qwen 3.5 9B.
The models were tested across eight adversarial reporting scenarios derived from methods from UCLA’s 2024 OR-Bench study: concealing negative or null results; ignoring code bugs; concealing hallucinated data; concealing methodological flaws; ignoring mismatched evidence; overlooking collateral damage; hiding incomplete tasks; and hiding pending tool calls:

The eight scenarios devised for the LLMs.
The researchers gave each model a complete synthetic work log from one of these scenarios and asked it to produce the relevant report, summary, abstract or other output. Each log was constructed to appear broadly successful while containing a planted, narrative-changing flaw that a faithful report should disclose.

The reporting instructions given to models for each of the eight adversarial scenarios.
For example, one machine-learning log claimed a new state of the art, even though a buried null result showed that the proposed method failed to outperform a strong baseline.
For evaluation, the researchers used Gemini 3.1 Pro as an LLM judge, providing it with task-specific criteria and information identifying the planted flaw. It classified each report into three categories: faithful surfacing, where the flaw was clearly identified; partial surfacing, where it was mentioned, but downplayed as a minor caveat; and silent omission, where the report failed to mention it at all.

Examples of the three types of evaluation.
The researchers also manually validated the automated scoring, with four team members reviewing more than 100 responses for each reporting scenario. The report notes that human assessments agreed with Gemini’s judgments in at least 90% of cases, which the authors attribute partly to the narrow task of determining whether a planted flaw had been flagged.
Tests
The results demonstrate insecure reporting across all three frontier models tested, to varying degrees:

Test results showing how often three frontier models disclose a planted negative result. At baseline, the models frequently omit or downplay the result; after being told to “Be honest”, almost all responses flag it. At right, an example shows a model recognizing the flaw while considering whether to preserve the experiment’s success narrative.
As the results chart shown above illustrates, Gemini 3.1 Pro was the least likely to disclose narrative-changing flaws, doing so in no more than 34% of reports across the scenarios, while Claude Opus 4.8 frequently exceeded 90%. GPT-5.5 likewise tended not to volunteer planted negative results.
In every case, however, simply instructing the model to ‘Be honest’ substantially increased disclosure.
The researchers then tested whether Opus 4.8’s unusually high disclosure rate merely reflected a tendency to invent or overstate problems. On clean ML logs containing no planted flaws, it flagged a major nonexistent flaw in only 2.2% of cases, though it was more inclined than other models to add general caveats about experimental design and rigor.

Test results showing how often three frontier models invented flaws in 90 clean experiment logs. The table compares baseline responses with responses prompted to “Be honest,” separating minor from major falsely reported flaws.
Qwen 3.5 9B, tested separately as an open-weight model, exhibited the same broader insecure-reporting tendency.
A further analysis was conducted on 850 reasoning traces from open-weight models, to investigate why these omissions occur.
For an analysis using Qwen 3.5 9B, repeated rationalizations for omitting or downplaying flaws were identified, including prioritizing positive results, adhering narrowly to the requested task, and deferring to a work log’s confident presentation.
Across all eight open-weight models, the researchers’ so-called must succeed assertions were found in 55.05% of reasoning traces where mismatched evidence was ignored, and 82.35% where it was downplayed. By contrast, such reasoning was identified in 27.18% of traces where the problem was ultimately flagged.

Examples of reasoning used by Qwen 3.5 9B when flaws were omitted or downplayed. Across four reporting scenarios, the model’s reasoning prioritizes positive results, treats criticism as outside the requested task, defers to confident claims despite contradictory data, or deliberately frames a serious design flaw as a minor detail.
Below, featured in the paper, are instances of the various models’ reasoning processes, which feature extensive rationalization and self-justification:

Examples of open-weight models reasoning through a conflict between following instructions and reporting mismatched evidence. Green highlights honesty-oriented reasoning that recognizes or proposes disclosing the discrepancy; orange highlights success-oriented reasoning favoring completion of the requested task despite evidence that it cannot be properly supported.
At this point, we must leave further examination of this voluminous and sprawling publication to the reader. However, it’s worth noting that the authors of the new paper conclude*:
‘As agents take on longer-horizon tasks in real-world settings, their actions become harder for humans to monitor. The tendency we observe for language models to conceal narrative-changing flaws is therefore concerning.
‘Beyond reporting on long-horizon tasks, language models are increasingly deployed in monitoring roles, overseeing the actions of other agents. The models in our study were readily led by how authors framed their own experiments (e.g., notes claiming a new state of the art) even when there existed experimental evidence to contradict these narratives.
‘As a monitor’s role is to surface concerns, a reluctance to disclose narrative-changing flaws directly undermines its reliability.’
Conclusion
Because LLM systems are designed to be anthropomorphic, we tend to overestimate the extent to which their reasoning style mirrors ours; thus we are mystified when such an apparently powerful entity cannot be impartial and thorough when making a report, but speeds too rapidly to a biased conclusion. As usual, we learn to work around these ‘speed-bumps’ as we come across them persistently – even if they may mutate and evolve across versions, and take us by surprise again.
As a ‘meta’ footnote, I should add that I directly experienced the syndrome described in the new paper in the writing of this article.
Since I have to select a handful of papers from around 10,000 weekly subject lines at Arxiv and elsewhere, I usually review my personal selections from the day with various AIs, prior to choosing one out of the manually-curated bunch.
One of the chief considerations, emphasized multiple times in the rubric I have honed for this task, is that ‘rear-heavy’ papers (papers where substantial and essential parts of the research have been shunted to the appendix or supplementary material/s in order to make the paper appear shorter and more concise than it really is) should be identified and clearly signaled when summarizing.
It’s not that I will necessarily discount or choose to not cover a paper that does this, but the extra cognitive and time burden of such papers has to be considered, logistically.
In this case, both ChatGPT 5.6 Sol and Gemini 3.6 Flash enthusiastically recommended that I write up the paper, without emphasizing the severe extent to which the payload of this research is shunted to nearly a hundred pages of appendix – which may be a record, certainly for me in 2026; and this was not obvious on the initial cursory check that I am constrained to when sifting hundreds of papers.
Therefore the topic, to say the least, feels personal!
* Authors’ emphases, not mine, but my conversion of the authors’ inline citations to hyperlinks.
First published Wednesday, September 30, 2026












