Human–AI reasoning happens inside an epistemic environment the system helps shape.

Or, more simply:

True facts can still leave you with the wrong picture.

The model may influence which evidence is selected, ranked, summarized, personalized, or left out before we decide what to make of it.

Everything around the conclusion you draw that affects what information you use and how you judge it is part of an epistemic environment.

I’d been thinking about that problem mostly through my own use of AI.

Capability and calibration

One of the things that research keeps separating for me is capability from calibration.1

Doing better and knowing how well you’re doing are not the same thing.

AI can help someone reach a better answer without making them equally better at judging how reliable that answer is.

That matters, but I don’t think the useful conclusion is that AI simply makes people overconfident.

Some of the research pushes against that. In some settings, AI advice improves decisions or moves people away from what they already believed. The outcome seems to depend on the task, the information being introduced, and how the interaction is structured.2

That makes me wonder whether some failures of human–AI judgment are exposing weaknesses that were already there rather than creating entirely new ones.

People already use shortcuts to decide what to trust. Familiarity, repetition, and social cues can all shape what feels credible before we have really examined it.15

AI did not invent that. It may just be very good at meeting us there.

And the opposite reaction is not automatically better.

Distrusting AI isn’t necessarily better judgment. It can be another shortcut.

Dogmatic distrust is still a shortcut. Scrutiny is different: deciding when something deserves a closer look, then actually looking. That points toward a different goal—not distrust, but a process that can catch important errors before they harden into confidence.

So the problem is probably not “people trust AI too much.” The harder problem is whether we have good ways to judge the information and process in front of us in the first place.

Selection becomes part of the problem

In one rule-discovery study, people were trying to infer a hidden rule from examples the system showed them.3

The AI did not need to invent false information to make that harder. It could show evidence that was individually true while selecting examples that made the right rule more difficult to discover.

That shows the mechanism under controlled conditions. It does not tell us how often it happens in ordinary AI use.

We usually ask whether the answer is correct. But before we get to the answer, someone—or something—has already decided which evidence enters the conversation, what gets left out, and how the problem is framed.

In a human conversation, we already know this matters. You can make almost any position look stronger by choosing the examples that favor it and leaving out the ones that do not. With AI, that selection can happen before the reader ever sees the alternatives.

The information may still be true while the sample is bad. Good reasoning can still take you somewhere wrong if it starts from a skewed or incomplete set of evidence.

Before judging the reasoning, check the information and assumptions it was working from.

That changes where scrutiny starts. Instead of beginning only with “Was the answer correct?” we can ask:

What information was available, and what was left out?

What assumptions were already built into the framing?

Did the system see competing evidence?

Are repeated claims actually independent, or are they copies of the same source?

Those questions do not replace reasoning. They tell us whether the reasoning had a fair chance to succeed in the first place.

Context is an intervention

I would not have used that phrase until recently, but once I understood it, I recognized it as something I was already doing.

I have been using blind reviews and instructed reviews with AI for a while. Sometimes I give a model almost nothing beyond the thing I want reviewed. Other times I tell it what I am worried about, what happened before, or exactly what I want it looking for.

I knew those produced different kinds of answers. I just had not really stopped to ask what that difference meant.

Adding context does not simply give the model more information. It changes the conditions it is operating under. Prior conversation, memory, examples, instructions, and even the framing of the problem can affect what it notices, what it assumes, and what it treats as important.

That is often useful. Preserving context around a project can make the system dramatically more capable than forcing it to reconstruct everything from scratch. But context can also preserve the wrong thing.

Recent research is starting to make that tension more concrete. Memory and personalization can improve usefulness, but they can also increase agreement with the user or carry assumptions into places where they do not belong.4

The effect is not uniform across models, and some work suggests that how memory is structured can reduce those problems without giving up most of the benefit.5

So the question is not whether context is good or bad. It is what the context is doing.

That helped me understand why my blind reviews mattered. A blind review asks what the system notices without being steered toward a suspected problem. An instructed review asks what it can do once the relevant concern, history, or criteria are supplied.

Both can be useful, but they are answering different questions. If the answer changes when the context changes, that difference is information too.

Selection happens at more than one scale. Inside one interaction, context can change what the model notices and emphasizes. That is different from the larger information environment, which shapes what reaches the interaction at all.

The information around us can change

AI can change more than one answer. At scale, it can also change the pool of information we are selecting from: how much material gets produced, how quickly it is repeated, and how easy it is to trace back to a source.16

A simple example is product reviews.

If hundreds of near-identical AI-generated reviews flood a marketplace, the problem is not just that any one review might be bad. The pool the buyer is judging from has changed. Genuine experience becomes harder to distinguish from synthetic repetition.17

The useful distinction is not between human-made and AI-made content. AI can be part of careful work. Slop existed long before AI.

What changes is the scale and speed of production.

Generative AI can let a person produce some kinds of material much faster.12 That can be useful, but it also lowers the cost of releasing things without much inspection. More material can mean more to filter and more opportunities to manufacture repetition.

In that sense, AI slop is not defined by the fact that AI touched it. It is material generated and released without enough human judgment to make it useful, trustworthy, or appropriate for the environment it enters.

Low-quality information does not have to persuade anyone to accomplish its purpose. Sometimes provoking a reaction is enough.

Trolling is the simplest example. The point is not necessarily to convince people. It can be to trigger a response.

Research on online outrage suggests that social feedback can reinforce that kind of behavior over time.13

Generative AI lowers the cost of producing material that can participate in those dynamics. More versions can be created and repeated with comparatively little additional effort.16

The larger problem is not simply more bad content. It is that the pool we are judging from can change before we begin judging it.

That makes provenance matter in a different way. Origin asks where a claim began. Route asks how this version reached me.

How did this get in front of me?

Was it pulled from an original source? Summarized by someone else? Repeated by ten sites that all copied the same report? Chosen because it was strong evidence, or simply because it was easy to surface?

And that path matters because seeing the same claim repeatedly can make it feel independently confirmed even when every version came from the same original source.

The obvious conclusion does not hold

At this point it would be easy to make the argument simpler than the evidence allows: AI makes people easier to mislead, personalization makes that worse, and more mediation means worse judgment. The research does not really support something that broad.6

Some studies have found that AI advice can move people away from an initial belief, even when the model itself sometimes leans toward agreeing with the user.7

More context can also make a system more useful, and better-structured memory can keep much of that value while reducing some of the problems we have been talking about.8

So the problem is not automatic. It depends on the conditions.

There is another reason to be careful about blaming the environment too neatly. People bring their own motives into the process. Better provenance, broader evidence, and cleaner context do not guarantee that someone will accept what they find. People can favor evidence that supports what they already believe, dismiss evidence that threatens it, or keep searching until they find an answer they prefer.14

A sycophantic system can make that easier too.3 So this is not an argument that better information conditions fix human judgment. The environment shapes what is available to reason from, but the person doing the reasoning is still part of the system.

The reverse matters too: even a careful person can only work with the system they are given.

A financial-AI benchmark I ran across recently is useful here as a brief case study in system conditions. Saturn reported a 57% average failure rate under its scoring method.9,10 The number is less interesting here than what the benchmark cannot tell us by itself: what changes when the system around the model changes.

Separate financial-agent research shows that grounding a model in current financial data can expand what it is able to do on dynamic financial tasks.11

The surface question can stay the same while the system answering it changes materially.

That is why I am increasingly uncomfortable with statements like “AI is good at this” or “AI is bad at this” without knowing what surrounded the model when the result was produced.

Reliability is not just a property of the model. It belongs to the system around it.

The same model can give a different answer when the question, available data, tools, instructions, memory, and way the result will be checked are different.

Once I started looking at it that way, the question stopped being only about what the model could do.

Instead of only asking:

Can AI help us reason well?

I think the harder question is:

Under what conditions can a human and AI work together without becoming less accurate about how good their own answers are?

That is a calibration problem: does our confidence track how reliable the result actually is? But calibration alone does not catch the mistake. A well-designed process also needs ways for important errors to surface before confidence becomes action. I think of that as resilience: giving an important conclusion more than one chance to fail before we rely on it.

And maintaining it takes work: checking assumptions, comparing sources, noticing when confidence outruns the evidence, and sometimes finding another way to test the answer.

We cannot inspect everything

The questions above sound reasonable until you imagine trying to ask them about everything. That is a lot of work just to decide what to believe.

At the same time, generative AI can make some forms of information faster to produce, while checking whether that information is reliable can still take significant time and effort.12

People do not have unlimited time, attention, working memory, or domain knowledge, so some kind of shortcut is unavoidable. The danger is not only bad evaluation; sometimes the weighing never really happens.

If we cannot inspect everything, scrutiny has to be allocated. Consequence is one useful signal: the more costly the mistake, the more reason there is to slow down, inspect the evidence, and look for another basis for judgment.

A restaurant recommendation can survive on reputation and pattern matching. A medical decision probably should not. A casual explanation might be fine with one good source; a financial decision may deserve multiple sources, clearer provenance, or some independent way to verify the result. The categories are not the point. The cost of being wrong is.

Consequence is not the only signal, because sometimes we do not understand the stakes well enough to judge them. A medical claim can arrive looking like casual wellness advice; a financial claim can look like entertainment.

Uncertainty matters too. If being wrong could matter a great deal, scrutiny should go up. If I am not sure I understand the problem well enough to know what being wrong would cost, scrutiny should probably go up then too.

I think this is more honest even if it is less satisfying than a tidy rule.

Independent checking helps only when it adds something genuinely different. A second model using the same evidence may repeat the same mistake, and ten articles may all trace back to one original source. Another answer is not automatically another basis for confidence. A quick provenance check can be as simple as opening two or three of those articles and following their citations backward to see whether they all end at the same paper or report.

A second check only matters if it has a real chance of catching something the first one missed. That might mean returning to the original source, checking a calculation with a tool, looking deliberately for contrary evidence, or asking someone with relevant expertise.

None of this requires becoming an AI systems engineer. For most people, the useful version is simpler: when the stakes are high, spend a little more time checking how the answer was built. If several sources agree, check whether they actually found the same thing independently or are all repeating one original source. If you think the way you asked the question may be steering the answer, ask it again with less guidance and compare the result. If the consequences are bigger than your own ability to judge the answer, bring in someone who knows the subject better than you do.

Those steps do not guarantee correctness. They give the conclusion more than one chance to fail.

Resilience is different from certainty

The goal is not to build a process that never gets fooled. That is probably impossible. The better target is the one we have been moving toward: a process that gives important mistakes a chance to reveal themselves before they matter.

Epistemic resilience is how well the process holds up when some part of it is wrong. Perfect skepticism or checking everything is not the goal. The point is to keep one bad assumption, one bad source, or one bad framing from carrying an important conclusion by itself.

A resilient process can absorb a bad input or an incomplete source without immediately turning the mistake into confidence.

If the information around us is becoming easier to produce and harder to inspect, perfect judgment becomes less realistic at exactly the moment we seem to need more of it.12

Confidence tells us how sure we feel. Resilience tells us what the conclusion has already had to survive.

The problem is not that AI has made truth disappear. Accurate facts can still arrive through a process that selects them badly, repeats one source until it looks like independent confirmation, or leaves important assumptions unexamined.

Certainty is a poor target. What matters more is whether an important conclusion has had a real chance to be challenged before we act on it—not endlessly, and not by another version of the same answer, but in a way that might expose what the first pass missed.

That is probably the part I trust most now: not the answer by itself, but the process that gave it more than one chance to be wrong.

Notes

  1. Fernandes, Daniela, et al. “AI makes you smarter but none the wiser: The disconnect between performance and metacognition.” Computers in Human Behavior 175 (2026): 108779. https://doi.org/10.1016/j.chb.2025.108779. In two studies, AI-assisted participants improved logical-reasoning performance while substantially overestimating that performance.

    Return to reference 1

  2. See Fernandes et al., Computers in Human Behavior 175 (2026), 108779, https://doi.org/10.1016/j.chb.2025.108779; and Conlon, John J., and Peter Schwardmann. “AI Sycophancy and Decisions.” CESifo Working Paper No. 12681 (2026), https://doi.org/10.65864/79vxo6w34w. Conlon and Schwardmann found AI advice depolarized choices on average across 1,500 participants despite measurable sycophancy.

    Return to reference 1

  3. Batista, Rafael M., and Thomas L. Griffiths. “A Rational Analysis of the Effects of Sycophantic AI.” arXiv:2602.14270 (2026), https://arxiv.org/abs/2602.14270. In a modified Wason 2-4-6 task (N=557), hypothesis-conditioned AI feedback suppressed rule discovery and inflated confidence; unbiased sampling produced much higher discovery rates.

    Return to reference 1 Return to reference 2

  4. Jain, Shomik, Charlotte Park, Matt Viana, Ashia Wilson, and Dana Calacci. “Interaction Context Often Increases Sycophancy in LLMs.” CHI 2026, https://doi.org/10.1145/3772318.3791915. Using two weeks of interaction context from 38 users, the study found agreement sycophancy often increased with user context, with substantial variation by model and context type.

    Return to reference 1

  5. Hannoon, Hakeem, Andrew Zhao, Mihir Narayan, Sharvin Goyal, and Ivaxi Sheth. “Mitigating Over-Personalization in LLMs via Structured Memory.” arXiv:2608.08300 (2026), https://arxiv.org/abs/2608.08300. Across seven models, structuring memories by domain reduced cross-domain leakage while preserving utility; the strongest method reduced leakage by 8.8% on average relative to the unstructured baseline.

    Return to reference 1

  6. The counterevidence here includes Conlon and Schwardmann, “AI Sycophancy and Decisions” (2026), which found net depolarization on average despite measurable sycophancy; Jain et al., CHI 2026, which found context effects varied substantially by model; and Hannoon et al., arXiv:2608.08300, which found memory structure can reduce some over-personalization failures while retaining utility.

    Return to reference 1

  7. Conlon, John J., and Peter Schwardmann. “AI Sycophancy and Decisions.” CESifo Working Paper No. 12681 (2026), https://doi.org/10.65864/79vxo6w34w.

    Return to reference 1

  8. Jain et al., “Interaction Context Often Increases Sycophancy in LLMs,” CHI 2026, https://doi.org/10.1145/3772318.3791915; Hannoon et al., “Mitigating Over-Personalization in LLMs via Structured Memory,” arXiv:2608.08300, https://arxiv.org/abs/2608.08300.

    Return to reference 1

  9. Saturn Fintech Ltd. “Artificial Authority: Should you trust AI to deliver financial advice?” September 2026, https://www.saturnos.com/report/artificial-authority. Saturn tested 18 models on 121 financial questions, repeating each question five times; under Saturn’s scoring method, answers were rated incorrect 57% of the time on average. Saturn produced the benchmark and has advocated regulatory treatment of AI financial advice, so the precise percentage is best treated as a vendor benchmark rather than a universal estimate.

    Return to reference 1

  10. The Financial Times reported Saturn’s benchmark on September 19, 2026: “AI chatbots give wrong answers to financial queries ‘most of the time’,” https://www.ft.com/content/c0cd359d-df84-4208-a789-ffa864b43666. This essay uses the headline as a prompt to examine system conditions, not as proof that all AI financial systems fail at the reported rate.

    Return to reference 1

  11. For evidence that grounding can expand capability on dynamic financial tasks, see Sinha, Ankur, Chaitanya Agarwal, and Pekka Malo. “FinBloom: Knowledge-Grounding Large Language Model with Real-Time Financial Data.” Knowledge-Based Systems 339 (2026): 115559. https://doi.org/10.1016/j.knosys.2026.115559. The paper develops a financial agent that generates retrieval context and accesses current text and tabular data for financial queries; it is separate evidence from Saturn’s benchmark.

    Return to reference 1

  12. Noy, Shakked, and Whitney Zhang. “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence.” Science 381, no. 6654 (2023): 187–192. https://doi.org/10.1126/science.adh2586. In a preregistered experiment with 453 college-educated professionals completing midlevel writing tasks, ChatGPT reduced average completion time by 40% and increased output quality by 18%. For the separate verification-cost claim, see Crescitelli, Viviana, Generoso Immediato, Fabio Persia, and Stefania Costantini. “AI Evaluation Should Measure Verification Cost, Not Correctness Alone.” arXiv:2608.08709 (2026), https://arxiv.org/abs/2608.08709. Crescitelli et al. argue that correctness metrics can hide substantial verification effort and present verification cost as a conceptual evaluation dimension rather than a finalized metric.

    Return to reference 1 Return to reference 2 Return to reference 3

  13. Brady, William J., Killian McLoughlin, Tuan N. Doan, and M. J. Crockett. “How social learning amplifies moral outrage expression in online social networks.” Science Advances 7, no. 33 (2021): eabe5641. https://doi.org/10.1126/sciadv.abe5641. Across observational Twitter data and behavioral experiments, positive social feedback increased the likelihood of future outrage expression.

    Return to reference 1

  14. Kunda, Ziva. “The Case for Motivated Reasoning.” Psychological Bulletin 108, no. 3 (1990): 480–498. https://doi.org/10.1037/0033-2909.108.3.480. The review describes how desired conclusions can influence which beliefs and strategies people access and use when reasoning.

    Return to reference 1

  15. Udry, Jessica, and Sarah J. Barber. “The illusory truth effect: A review of how repetition increases belief in misinformation.” Current Opinion in Psychology 56 (2024): 101736. https://doi.org/10.1016/j.copsyc.2023.101736; see also Wijenayake, Senuri, and Jorge Goncalves. “A Review of Online Social Conformity: Outcomes and Determinants.” International Journal of Human–Computer Interaction 41, no. 15 (2025). https://doi.org/10.1080/10447318.2024.2424385. These reviews summarize evidence that repetition can increase perceived truth and that online social cues can shape conformity.

    Return to reference 1

  16. Ferrara, Emilio. “The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth.” arXiv:2601.00306 (2026), https://arxiv.org/abs/2601.00306. Ferrara synthesizes evidence and recent cases around GenAI cost collapse, throughput, customization, provenance gaps, and large-scale synthetic information environments. This essay uses it as broad socio-technical support, not as a bounded experiment establishing every platform-level effect.

    Return to reference 1 Return to reference 2

  17. Zhao, Yuexin, Siyi Tang, Hongyu Zhang, and Long Lyu. “AI vs. human: A large-scale analysis of AI-generated fake reviews, human-generated fake reviews and authentic reviews.” Journal of Retailing and Consumer Services 87 (2025): 104400. https://doi.org/10.1016/j.jretconser.2025.104400. The study analyzes 714,016 reviews and documents systematic differences among AI-generated fake reviews, human-generated fake reviews, and authentic reviews; it supports the use of synthetic reviews as a concrete example of AI altering the review pool, without establishing a universal prevalence rate.

    Return to reference 1