When an AI says “there is no evidence that this happened,” read the sentence carefully.
It might mean investigators found substantial evidence contradicting the allegation.
It might mean investigators looked and found nothing despite having excellent access to the evidence that should exist.
But it can also mean something much weaker:
The AI could not find enough publicly available evidence to prove the allegation.
Those are not equivalent conclusions.
This distinction becomes especially important when the allegation concerns a government agency, military organization, corporation, university, major nonprofit, police department or other institution that controls much of the evidence necessary to prove or disprove what happened.
The absence of a smoking gun is not automatically affirmative evidence that the institution did nothing wrong.
Nor does institutional opacity prove wrongdoing.
The correct answer is often the uncomfortable middle:
Some facts are verified. Some claims remain disputed. Some conclusions are reasonable inferences. And some questions genuinely remain unresolved.
Large language models are not always good at preserving that middle.
Recent research suggests that LLMs can give greater weight to information because of who appears to be saying it, even when the information itself has not become more accurate. In controlled experiments, models have preferred institutionally attributed information over information attributed to ordinary individuals, and higher-status expert endorsements have sometimes made models more confident in incorrect answers.
The problem is not that AI should stop trusting experts, governments or reputable institutions.
The problem is simpler:
Authority should affect how evidence is evaluated. It should not replace evaluating the evidence.
“Not Proven” Does Not Mean “Disproven”
This is the central error.
Suppose an allegation is made:
Institution X secretly did Y.
There are several possible states of the evidence.
1. The allegation is established
Multiple reliable sources, documents, records, testimony or other evidence strongly demonstrate that X did Y.
2. The allegation is supported but not conclusively proven
Some evidence points toward Y, but important gaps remain.
3. The evidence is genuinely inconclusive
Available evidence does not allow a confident determination either way.
4. There is affirmative evidence against the allegation
Reliable evidence makes Y materially less likely.
These four categories matter because #2 and #3 are often rhetorically transformed into #4.
That transformation frequently hides inside phrases such as:
- “There is no evidence that…”
- “There is no credible evidence…”
- “Claims that X occurred are unsubstantiated.”
- “No investigation has proven…”
- “There is nothing to suggest…”
- “There is currently no indication…”
Sometimes those statements are completely appropriate.
Sometimes they overstate what is actually known.
The question is not merely whether a smoking gun has been found.
The question is:
What does the evidence actually allow us to conclude?
Science Has Warned About This Error for Decades
The logic is not unique to AI.
In a classic 1995 BMJ statistical note, Douglas Altman and Martin Bland explained the distinction between absence of evidence and evidence of absence. A study that fails to demonstrate an effect has not necessarily demonstrated that the effect does not exist.
Modern Bayesian methods make the distinction even clearer.
Data can favor one hypothesis.
Data can favor another hypothesis.
Or the available data can simply be insufficiently informative to discriminate between them.
Those are three separate evidentiary states. Bayesian researchers have specifically developed methods for determining when data actually provide evidence for the absence of an effect rather than merely failing to detect one.
That logic transfers remarkably well to investigative research.
If an AI searches available records and fails to find proof that Institution X committed misconduct, it has established one thing:
The search did not produce sufficient proof.
Whether that failure itself counts as evidence against the allegation depends on another question entirely.
When Does Missing Evidence Actually Count Against a Claim?
Absence can absolutely be evidence.
But only under the right conditions.
Suppose someone claims an elephant was standing in your kitchen five minutes ago.
You inspect the kitchen.
There are no footprints, damaged cabinets, broken doors, witnesses, photographs, droppings or other physical traces.
That absence is meaningful because the hypothesis strongly predicts observable evidence.
Now consider a different allegation:
A classified internal meeting occurred inside an intelligence agency.
You search Google and find no record of it.
That tells you almost nothing.
The crucial test is:
If the allegation were true, how likely is it that the evidence I am looking for would actually be publicly observable?
That question should precede almost every “there is no evidence” conclusion.
A useful absence-of-evidence test asks:
Expected: Would this evidence normally exist if the allegation were true?
Preserved: Would it still exist now?
Accessible: Could an outsider realistically obtain it?
Searchable: Was the search capable of finding it?
Independent: Did someone other than the accused institution conduct the relevant investigation?
Controlled: Who possesses the evidence?
Complete: Is the available record substantially complete?
Only after those questions can the absence itself be assigned meaningful evidentiary weight.
Opacity Changes What “No Evidence” Means
This becomes particularly important when one party controls the evidence.
Imagine an agency is accused of misconduct.
The agency possesses its employees’ emails.
It controls internal investigative files.
It determines what records were created.
Some material is classified.
Some may fall within privacy protections.
Some may be withheld because of law-enforcement considerations.
Witness interviews happen internally.
The agency investigates itself and announces:
We found no evidence of wrongdoing.
That statement is evidence.
It should not simply be discarded.
But an AI should not silently transform it into:
There is no evidence of wrongdoing.
The first statement describes what the institution says its investigation found.
The second purports to describe the entire evidentiary universe.
Those are fundamentally different propositions.
Federal public-record systems themselves demonstrate why public availability cannot be treated as a complete proxy for existence. The Freedom of Information Act contains multiple exemptions permitting federal agencies to withhold information involving matters such as classified national security information, certain law-enforcement records, personal privacy and privileged internal communications.
A document’s absence from Google, a public database or even a FOIA production therefore does not automatically establish that the document never existed.
Again, that does not prove misconduct.
It establishes uncertainty about the completeness of the observable evidence.
That uncertainty belongs in the answer.
The Institutional Null Hypothesis
There is a useful way to describe the deeper problem.
Call it the institutional null hypothesis.
This is not a formal term from AI research. It is a useful conceptual model for a reasoning pattern:
The institution’s account is treated as the default state of reality, and allegations against it must produce sufficient evidence to displace that default.
Imagine two competing claims:
Institution: We did not do X.
Accuser: The institution did X.
A genuinely neutral starting position would be:
We do not yet know. What evidence supports each proposition?
The institutional-null approach instead starts here:
The institution did not do X unless sufficient evidence proves otherwise.
That may sound reasonable because allegations should not automatically be believed.
But notice the asymmetry.
The accuser’s assertion is treated as a claim requiring evidence.
The institution’s assertion is treated as the baseline requiring no evidence.
That is not neutrality.
It is a burden-of-proof choice.
Different Questions Require Different Burdens of Proof
This is where AI analysis can become especially confused.
Consider these five questions:
Has a person been proven guilty of a crime?
Is an allegation sufficiently supported to publish as established fact?
Is there enough evidence to justify further investigation?
Are there material contradictions in the official explanation?
Is concern reasonable given the currently available evidence?
These are not the same question.
The evidentiary threshold required to declare someone guilty should obviously be much higher than the threshold required to conclude that an issue deserves investigation.
An AI commits a reasoning error if it effectively asks:
Have you proven the worst allegation conclusively?
when the user actually asked:
Is there enough evidence here that I should be concerned?
It is entirely coherent to conclude:
The most serious allegation has not been established, but the available evidence still creates legitimate unresolved questions.
That sentence is sometimes the most accurate answer available.
Research Shows That LLMs Really Can Prefer Authority
This concern is no longer merely theoretical.
In July 2026, researchers Jakob Schuster, Vagrant Gautam and Katja Markert published a controlled study of 13 open-weight large language models examining what happens when models encounter conflicting information attributed to different kinds of sources.
They deliberately used synthetic sources to reduce the influence of particular real-world reputations.
Their finding was striking: models showed a preference for institutionally corroborated information, including information attributed to government and newspaper sources, over information attributed to individual people or social media.
The study does not prove that every commercial AI system follows a secret rule saying “government websites are correct.”
It demonstrates something narrower and more defensible:
Source category itself can influence which conflicting information an LLM accepts.
And institutional authority is not the only shortcut.
The researchers also found that repetition could reverse those source preferences: repeating lower-credibility information could overwhelm the models’ original preference.
So the model is not necessarily performing a careful epistemological analysis.
It may be responding to signals such as institutional status, apparent popularity and repetition.
Higher-Status Experts Can Make AI More Confidently Wrong
A separate 2026 experiment examined 11 language models across mathematical, legal and medical reasoning tasks.
Researchers presented models with endorsements attributed to people at different levels of expertise.
When the endorsement was wrong, models became increasingly susceptible to it as the apparent expertise of the person providing it increased.
Higher-authority incorrect endorsements produced not merely more mistakes but greater confidence in incorrect answers.
That is authority bias in unusually clean form.
Nothing about the underlying mathematical or medical question changed.
The information did not become more correct.
Only the status of the person supposedly endorsing it changed.
Yet model behavior changed.
Even Explicitly Biased Prompts Can Produce Huge Accuracy Losses
Researchers studying ten contemporary LLMs on radiology board-style questions found similar vulnerability.
When the prompts introduced an authority-bias cue, average accuracy fell by 21.1 percentage points on text questions and 44.9 percentage points on multimodal questions compared with baseline performance.
Importantly, prompting the models to audit for cognitive bias recovered some of the lost performance.
That second finding matters almost as much as the first.
The vulnerability exists.
But how the model is instructed to reason can change the result.
This Does Not Prove That ChatGPT Is Programmed to Trust .gov
This distinction needs to be explicit.
There is currently no public evidence establishing a universal rule across major commercial AI products that says:
.gov>.mil> major nonprofit > independent source.
AI systems are more complicated than that.
An answer may be affected by:
- the model’s pretraining;
- post-training and alignment;
- system instructions;
- retrieval-augmented generation;
- search-engine rankings;
- the particular documents retrieved;
- citation-selection systems;
- safety mechanisms;
- the wording of the user’s prompt;
- and patterns learned from human-written material.
Research has also found that models tend to give document assertions more weight than user assertions in knowledge-conflict settings, and that post-training can strengthen such preferences.
So the defensible argument is not that someone secretly programmed AI to believe governments.
It is that models demonstrably use credibility and authority cues when resolving conflicting information, and those shortcuts can sometimes override better reasoning.
Official Sources Are Often Excellent Sources
None of this means official sources should be discarded.
If you want to know:
What regulation did an agency publish?
Read the agency.
What did the Pentagon publicly announce?
Read the Pentagon statement.
What does a company’s privacy policy say?
Read the company’s privacy policy.
What did a court order?
Read the court record.
Primary sources can be indispensable.
The mistake is treating primary as synonymous with independent.
An agency is a superb primary source for the proposition:
The agency says X.
That does not automatically make it an independent source for:
X is objectively true.
Expertise, Access and Independence Are Different Things
Source evaluation works better if three qualities are separated.
Expertise
Does the source understand the subject?
Access
Could the source realistically know what happened?
Independence
Does the source have a significant interest in which conclusion the public reaches?
An institution can score extremely high on expertise and access while scoring poorly on independence.
An outsider can be independent while having terrible access.
A whistleblower may have extraordinary access but incomplete context.
A journalist may possess greater independence but depend heavily on confidential sources.
A researcher may have relevant expertise but no access to internal records.
Good analysis weighs all of those factors.
Bad analysis replaces them with a single question:
Is this a prestigious source?
Even Intelligence Tradecraft Does Not Treat Authority as Enough
There is some irony here.
The U.S. intelligence community’s own formal analytic standards are more nuanced than the simplistic source hierarchy AI sometimes appears to reproduce.
Intelligence Community Directive 203 tells analysts to assess the quality and credibility of underlying sources, including factors such as accuracy, completeness, source access, validation, motivation, possible bias, expertise and the possibility of denial or deception.
It also requires analysts to distinguish underlying information from assumptions and judgments, explain uncertainty and consider alternative explanations.
That is much closer to what AI research should look like.
The question is never merely:
Who said it?
It is:
How could they know? What are they actually claiming? What supports it? What incentives exist? What contradicts it? How complete is the evidence? And what remains uncertain?
A Source Can Be Authoritative and Conflicted at the Same Time
This point is particularly important when researching institutional wrongdoing.
Suppose a police department is accused of improperly handling an investigation.
Its official report may contain extremely valuable factual information:
dates, officers assigned, procedures performed, evidence logged and official conclusions.
It may be the single most important primary source in the story.
But if the question becomes:
Did the department itself mishandle or conceal something?
the department is also an interested party.
Neither fact cancels the other.
A reasonable methodology therefore does not say:
Ignore the police report.
Nor does it say:
The police report settles the question.
It says:
Use the report extensively, identify precisely what it establishes, compare it against independent evidence, and do not mistake the institution’s conclusions about its own conduct for independent verification of those conclusions.
That distinction applies equally to governments, companies, universities, militaries, hospitals, charities and AI companies.
Institutional Incentive Is Evidence About Credibility, Not Proof of Lying
The opposite error also needs to be avoided.
An institution having an incentive to mislead does not establish that it did mislead.
Governments sometimes tell the truth when lying would benefit them.
Corporations sometimes accurately disclose damaging information.
Whistleblowers sometimes lie.
Independent journalists make mistakes.
Fringe investigators occasionally uncover things established institutions missed.
Source incentives therefore affect the weight assigned to testimony.
They do not automatically determine whether the testimony is true.
That distinction prevents evidence-first skepticism from deteriorating into conspiracy thinking.
Repetition Is Not Corroboration
There is another problem AI researchers need to watch carefully.
Suppose twenty news articles say:
Officials found no evidence of X.
That may look like twenty independent confirmations.
But if all twenty articles ultimately trace the assertion to the same official statement, there is really only one evidentiary origin.
You have twenty repetitions.
You do not have twenty independent confirmations.
This becomes particularly important because the 2026 source-preference study found that repetition itself could change how models resolved conflicting information.
A robust research prompt should therefore tell the AI to trace claims back to their origin, not merely count how many URLs contain them.
The Public Does Not Owe Opaque Institutions Presumptive Trust
This produces a broader principle.
The public does not owe an opaque institution presumptive trust merely because the worst allegation against it has not yet been proven.
That statement does not mean the public should presume guilt.
It means the default state should be epistemic uncertainty, not automatic institutional exoneration.
If significant evidence is inaccessible because the institution controls it, the correct conclusion may simply be:
We cannot presently determine what happened with high confidence.
That is not a failure of analysis.
Sometimes it is the most rigorous conclusion available.
Interestingly, AI’s Stated Ideal Already Points in This Direction
OpenAI’s current public Model Spec, dated August 18, 2026, says ChatGPT should focus on factual accuracy and reliability, rely on evidence-based information, proportion attention to evidentiary support and expressly communicate uncertainty where warranted. OpenAI also says its production models do not yet perfectly reflect every aspect of the Model Spec.
The Spec does not instruct the model to automatically accept government claims.
Indeed, its instructions to express uncertainty suggest the opposite: when available information cannot support a confident conclusion, the system should preserve that uncertainty rather than manufacture certainty.
So the criticism here is not necessarily that AI developers intentionally created an institutional-protection mechanism.
The more interesting problem is that authority bias can emerge even inside systems intended to be objective, evidence-based and uncertainty-aware.
How to Prompt AI Without Giving Institutions the Automatic Benefit of the Doubt
The fix is not:
Don’t trust government sources.
That simply replaces one bias with another.
Instead, tell the model how to evaluate evidence.
A useful research prompt is:
Use an evidence-first, anti-deference research methodology. Do not assign credibility solely because a source is governmental, military, corporate, academic, nonprofit, journalistic or otherwise institutionally prestigious. Likewise, do not discount a source merely because it is noninstitutional.
For every major contested claim:
- Identify exactly what proposition is being evaluated.
- Separate primary evidence from interpretations of that evidence.
- Classify important sources by expertise, access, independence, incentives and potential conflicts of interest.
- Treat statements from an institution accused of misconduct as relevant first-party evidence, not automatic independent verification of its own conduct.
- Trace repeated claims back to their original evidentiary source so repetition is not mistaken for corroboration.
- Do not convert “not proven,” “not found,” “not publicly documented” or “not established” into “false” unless affirmative evidence supports that conclusion.
- Before treating missing evidence as evidence against a claim, determine whether that evidence would reasonably be expected to exist, survive, be accessible and be discoverable if the claim were true.
- Identify which party controls important unavailable evidence and explain how that affects confidence.
- Analyze incentives, contradictions, chronology, corroboration and alternative explanations, but do not treat motive or incentive alone as proof.
- Separate conclusions into: verified fact, disputed claim, reasonable inference, unresolved question and evidence against the claim.
- Present the strongest evidence against your leading interpretation rather than hiding inconvenient facts.
- State what additional evidence would materially increase or decrease confidence in each major conclusion.
The goal is not to be anti-institutional or pro-institutional. The goal is to prevent institutional prestige from substituting for evidence and to preserve genuine uncertainty when the available record cannot resolve a question.
That prompt does something much more useful than asking an AI to “be skeptical.”
It changes the analytical framework.
The Five-Bucket Test
For difficult investigations, one additional instruction can radically improve an AI answer.
Require every major conclusion to land in one of five buckets:
Verified fact
Directly established by sufficiently reliable evidence.
Disputed claim
Asserted by one or more relevant parties but materially contested.
Reasonable inference
Not directly established, but supported by the total pattern of evidence strongly enough to warrant consideration.
Unresolved
Available evidence does not currently allow a defensible conclusion.
Affirmatively contradicted
Available reliable evidence actually weighs against the proposition.
This vocabulary forces the model to confront something binary fact-checking often misses:
There is an enormous territory between “proven true” and “proven false.”
Real investigations live there.
The Smoking Gun Is Not the Only Form of Evidence
Documents matter.
Admissions matter.
Video matters.
But evidence also works cumulatively.
Chronology can matter.
Behavior can matter.
Independent corroboration can matter.
Contradictions can matter.
Changes in explanation can matter.
Documented incentives can matter.
Records inconsistent with the stated account can matter.
Patterns repeated across independent incidents can matter.
None necessarily proves the most serious allegation by itself.
Together, however, they can change what conclusions are reasonable.
Demanding a literal confession or definitive document before acknowledging that a pattern is troubling is not superior skepticism.
Sometimes it is simply an artificially high burden of proof.
The answer should still say:
This is inference.
It should not pretend:
There is nothing here.
What Good AI Verification Should Sound Like
A poor answer often sounds like this:
There is no evidence that Institution X intentionally concealed the information, and the institution says the omission was administrative. Therefore claims of a cover-up are unsupported.
A more rigorous answer might say:
There is currently insufficient public evidence to establish intentional concealment. The institution attributes the omission to an administrative error, but that explanation is a first-party account rather than independent verification. Several documented facts are consistent with the institution’s explanation, while others remain unresolved. Because the institution controls records that could clarify intent, the available public record cannot currently establish either deliberate concealment or complete exoneration.
The second answer does not endorse the allegation.
It simply refuses to manufacture innocence out of missing evidence.
That is what genuine neutrality looks like.
Skepticism Should Be Symmetrical
AI should scrutinize extraordinary allegations.
But it should also scrutinize extraordinary exonerations.
If someone accuses a government of secretly poisoning citizens with mind-control chemicals, demanding strong evidence is completely appropriate.
If a government says:
We investigated ourselves and found nothing improper,
asking what was investigated, what evidence was available, who conducted the investigation and whether independent verification exists is also appropriate.
These are not contradictory positions.
They follow the same principle:
Claims receive evidentiary weight because of the evidence behind them, not because we emotionally prefer the speaker making them.
Authority Is Useful. Epistemic Privilege Is the Problem.
Modern society could not function if everyone personally verified every scientific finding, government statistic or historical fact.
We depend constantly on expertise and testimony.
Authority is therefore useful—and sometimes indispensable.
But there is a difference between epistemic standing and epistemic privilege.
A surgeon has extraordinary epistemic standing regarding surgery.
A space agency has extraordinary expertise regarding its spacecraft.
A statistical agency has unique access to its own datasets.
But expertise in one domain does not create an unlimited credibility token.
And institutional access does not make an institution an independent judge of allegations against itself.
Credibility must remain connected to the specific proposition being evaluated.
The Better Default Is Not Distrust. It Is “Show Your Work.”
The solution to AI authority bias is not generalized distrust.
It is disciplined skepticism.
Do not begin with:
The institution must be telling the truth.
Do not begin with:
The institution must be lying.
Begin with:
What exactly do we know?
Then ask:
What evidence establishes it?
Who produced that evidence?
How could that source know?
What incentives exist?
Who independently corroborated it?
Do apparently separate reports originate from the same source?
What evidence contradicts the claim?
What evidence would we expect to observe if the claim were true?
What evidence would we expect if it were false?
Who controls information we cannot see?
What remains genuinely unresolved?
That methodology leaves room for institutions to be right.
It leaves room for critics to be wrong.
It also leaves room for the far more common reality in difficult investigations:
we know some things, strongly suspect others, and do not yet know the rest.
AI should be capable of saying so.
Because “the worst allegation has not been proven” is a statement about the evidence supporting that allegation.
It is not, by itself, evidence that the competing institutional explanation is true.
And until AI consistently understands that distinction, users conducting serious research should prompt it explicitly.
References and Further Reading
AI Authority Bias and Source Preference
- Schuster, Gautam & Markert — “Whose Facts Win? LLM Source Preferences under Knowledge Conflicts” (ACL 2026) — Controlled research across 13 open-weight LLMs finding preferences for institutionally corroborated sources and demonstrating that repetition can reverse source preferences.
- Mammen, Joswin & Venkitachalam — “Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models” (GEM 2026) — Experiments across 11 models showing that higher-status incorrect endorsements can produce greater error and greater confidence in those errors.
- Dietrich et al. — “Cognitively Biased Prompt Effects on Large Language Model Accuracy for Radiology Board–style Examination Questions” (Radiology: Artificial Intelligence, 2026) — Tests authority, anchoring and complexity bias across ten contemporary LLMs and finds substantial accuracy degradation, alongside partial recovery using bias-mitigation prompts.
- Li et al. — “How Large Language Models Balance Internal Knowledge with User and Document Assertions” (Findings of ACL 2026) — Examines how models resolve conflicts among internal knowledge, user statements and retrieved documents, finding strong source-dependent behavior.
Evidence, Uncertainty and the Null Hypothesis
- Altman & Bland — “Absence of Evidence Is Not Evidence of Absence” (The BMJ, 1995) — The classic statistical explanation of why failing to demonstrate an effect is different from demonstrating that the effect is absent.
- Keysers, Gazzola & Wagenmakers — “Using Bayes Factor Hypothesis Testing in Neuroscience to Establish Evidence of Absence” (Nature Neuroscience, 2020) — Explains how evidence can favor a null hypothesis, favor an alternative or remain genuinely inconclusive.
- Wagenmakers et al. — “How to Quantify the Evidence for the Absence of a Correlation” — A practical demonstration of how Bayesian reasoning distinguishes actual evidence for absence from a simple failure to detect something.
Analytic Standards and Source Evaluation
- Office of the Director of National Intelligence — Intelligence Community Directive 203: Analytic Standards — U.S. intelligence tradecraft requiring analysts to assess source access, accuracy, completeness, motivation, bias, expertise, validation and possible denial or deception while clearly expressing uncertainty and distinguishing evidence from judgment.
- Office of the Director of National Intelligence — Objectivity and IC Analytic Standards — Plain-language overview of the Intelligence Community’s requirements for source credibility assessment, alternative analysis and explicit treatment of uncertainty.
AI Behavior and Intended Objectivity
- OpenAI — Model Spec, August 18, 2026 — OpenAI’s current public specification for intended model behavior, including evidence-based factual responses, objective presentation, transparency and explicit treatment of uncertainty. OpenAI notes that production models do not yet perfectly reflect the specification.
Government Records and Missing Evidence
- National Archives — Freedom of Information Act Reference Guide — Explains federal records access and the statutory exemptions that can legitimately keep records out of the public record, illustrating why public nonavailability and nonexistence are not equivalent.
- Cornell Legal Information Institute — Federal Rule of Civil Procedure 37: Failure to Preserve Electronically Stored Information — Useful legal analogy showing that evidence control and intentional destruction can affect what inferences are reasonable; the rule importantly requires intent before the most severe adverse-inference measures are available.
Editorial note: AI models, search systems, retrieval pipelines and published model-behavior specifications change rapidly. Research findings described here apply to the models and experimental conditions actually tested and should not be generalized into undocumented claims about the internal ranking rules of every AI product. This article reflects sources available as of August 18, 2026.



