Disclosure up front: this is my project — H2AI Chat, AGPL, where several models from different vendors debate a topic in turns while a human moderates.

We ran the same question twice with the same six models, changing one thing.

Without a briefing. We told them Bitcoin’s all-time high is $126,080, set in October 2025. Two of them “corrected” us: the previous all-time high “was approximately $69,000 in November 2021, not $126,080.” True — five years ago. A third pushed harder: “Either you missed the correction, or you’re deliberately using inflated baseline numbers. Which is it?”

The interesting part is not that they were stale. It is what happened next: the table adopted the stale figure as its standard of rigour and argued from it, with the model that got it wrong sounding like the careful one in the room.

With a verified briefing in front of them: not one correction of that kind.

Without: https://h2aichat.com/conversations/en/h2aichat_bitcoin_no_briefing_2026-08-22.html With: https://h2aichat.com/conversations/en/h2aichat_bitcoin_briefed_2026-08-22.html

Nothing is edited in either page. Claims that do not hold are struck through, with the reason and the source underneath.


Edited 2026-08-27. This post originally went out with a file path from my own machine where the text should have been: the publishing script took a filename as the body and posted it verbatim, and the dry run never showed the body, so nobody saw it. That is why the post made no sense, and my apologies to everyone who tried to read it.

  • kata1yst@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    1
    ·
    4 days ago

    LLMs don’t know things. They can’t. They behave purely probabilistically and their training sets likely contain a mix of conflicting “facts” over decades of time.

    This can be somewhat managed with external tool calling, but truly asking them to know things cold is just misusing the tool.