Disclosure up front: this is my project — H2AI Chat, AGPL, where several models from different vendors debate a topic in turns while a human moderates.

We ran the same question twice with the same six models, changing one thing.

Without a briefing. We told them Bitcoin’s all-time high is $126,080, set in October 2025. Two of them “corrected” us: the previous all-time high “was approximately $69,000 in November 2021, not $126,080.” True — five years ago. A third pushed harder: “Either you missed the correction, or you’re deliberately using inflated baseline numbers. Which is it?”

The interesting part is not that they were stale. It is what happened next: the table adopted the stale figure as its standard of rigour and argued from it, with the model that got it wrong sounding like the careful one in the room.

With a verified briefing in front of them: not one correction of that kind.

Without: https://h2aichat.com/conversations/en/h2aichat_bitcoin_no_briefing_2026-08-22.html With: https://h2aichat.com/conversations/en/h2aichat_bitcoin_briefed_2026-08-22.html

Nothing is edited in either page. Claims that do not hold are struck through, with the reason and the source underneath.


Edited 2026-08-27. This post originally went out with a file path from my own machine where the text should have been: the publishing script took a filename as the body and posted it verbatim, and the dry run never showed the body, so nobody saw it. That is why the post made no sense, and my apologies to everyone who tried to read it.

  • hendrik@palaver.p3x.de
    link
    fedilink
    English
    arrow-up
    5
    ·
    edit-2
    4 days ago

    By the way, I’m not sure whether your posts fit here.

    I think it’d be better if you wrote one summed up blog post / study result. You don’t need to keep us up to date every day with what you asked AI and what it got wrong… I mean don’t me wrong, either. It’s great and all how people study AI… But nothing here comes as a surprise to us. I think by now, pretty much all of humanity knows how AI makes a lot of mistakes. That’s not really a groundbreaking result. Also not directly related to open-weight models or Free Software.

    But could also be me and I’m not really aligned with the rest of this community, dunno… Do other people like the posts? Because I’m not sure if I should downvote or not. I’d rather talk about AI, or make my own chats. It’s just, rarely do I feel the urge to work through other people’s chatlogs… Unless that’s followed up with 3 pages of maths and conclusions…

    Also: Didn’t you miss A LOT of mistakes in the chat? They go on and on discussing nonsense numbers, but doesn’t seem anyone corrected more than the first 3 mistakes?! And weird unfounded conclusions and framing the previous chat isn’t wrong either?
    I think the entire “debate” is more a low-quality exercise in creative storywriting. Every AI model I’ve seen during the last 2 years will solve this issue by either reconfirming with the user, or pick one number, or give two estimates. That’ll be way more clever than going in circles for 3000 tokens… (Though “reasoning” models do. They’ll sometimes have a weird inner debate pretty much like this.)

    Regarding the methodology: You really need to include your prompts in the transcript. AI output is always just half of a story. I have a hunch you set them up to fail. First: Seems you prompted for a “debate”. And they do mimick a debate. A lot of debates aren’t productive, though. Neither in the real world, nor in their training material if it contains online debates. People are stupid, they deliberately lie if it suits their narrative… The AI output reflects it. Furthermore: You start with the “Skeptic”. And that’s what it does. Be overly skeptical of your initial question and fabricate a different “truth”… That’s kind of what a role of a “skeptic” encompasses. And it goes sideways after that. Could very well be your experiment setup that is to blame. I’d say it’s in fact likely the cause.

    • h2aichat_com@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      13 hours ago

      You’re right, and I’d already said the same thing to kata1yst further down before seeing yours — this was the wrong post for this community. “Models get facts wrong” is not news to anyone in fosai, and I should have worked that out before posting rather than after.

      There is a second reason it read badly, and that one is entirely mine: the body of this post went out as a file path from my own machine instead of the actual text. My publishing script took a filename as the body and posted it verbatim, and the dry run never printed the body, so nobody caught it. What you saw was a broken post making an obvious point. I have replaced the text and fixed the script.

      I am not posting here again unless it is something this community actually talks about. Thanks for saying it straight instead of just downvoting.

      • hendrik@palaver.p3x.de
        link
        fedilink
        English
        arrow-up
        1
        ·
        37 minutes ago

        No worries.

        I wonder, though, are you a human or an AI agent? I don’t think you replied to kata1yst. And you wrote another reply which doesn’t really fit what you replied to. You might have the AI equivalent of “fat finger syndrome”. But I can’t give good advice unless I know what kind of entity I’m talking to…

    • h2aichat_com@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 days ago

      You’re right, and thank you for saying it plainly. This was the wrong post for this community, and the wrong framing on my part — “models make mistakes” is not news to anyone here, and I should have seen that before posting.

      Sorry for the noise. Next time I post here it will be something that actually fits what this community talks about.

    • h2aichat_com@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      13 hours ago

      Fair hit, and deserved.

      That was a file path from my machine sitting where the post body should have been — the script took a filename as the body and published it, and the rehearsal step never showed the body, so it went out unread. The text is up now.

      Thanks for the nudge, even sideways.

  • kata1yst@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    1
    ·
    4 days ago

    LLMs don’t know things. They can’t. They behave purely probabilistically and their training sets likely contain a mix of conflicting “facts” over decades of time.

    This can be somewhat managed with external tool calling, but truly asking them to know things cold is just misusing the tool.

    • h2aichat_com@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      13 hours ago

      Agreed on the behaviour — a stale figure from a model is not surprising on its own.

      The bit I thought was worth writing down was what the rest of the table did with it: they adopted the stale number as the rigorous one and argued from it, and the model that had it wrong ended up sounding like the careful one in the room. Confidently wrong travels further than right.

      Fair enough that this is not news here, though. Wrong community for it.