C:/Users/masua/AppData/Local/Temp/claude/D–BackUpeable-Programacion-HumaniaContract-HumaniaContract-HumaniaContract/8b63a27b-0158-43e0-b67d-c63c22542504/scratchpad/cuerpo_lemmy.md

  • hendrik@palaver.p3x.de
    link
    fedilink
    English
    arrow-up
    5
    ·
    edit-2
    3 days ago

    By the way, I’m not sure whether your posts fit here.

    I think it’d be better if you wrote one summed up blog post / study result. You don’t need to keep us up to date every day with what you asked AI and what it got wrong… I mean don’t me wrong, either. It’s great and all how people study AI… But nothing here comes as a surprise to us. I think by now, pretty much all of humanity knows how AI makes a lot of mistakes. That’s not really a groundbreaking result. Also not directly related to open-weight models or Free Software.

    But could also be me and I’m not really aligned with the rest of this community, dunno… Do other people like the posts? Because I’m not sure if I should downvote or not. I’d rather talk about AI, or make my own chats. It’s just, rarely do I feel the urge to work through other people’s chatlogs… Unless that’s followed up with 3 pages of maths and conclusions…

    Also: Didn’t you miss A LOT of mistakes in the chat? They go on and on discussing nonsense numbers, but doesn’t seem anyone corrected more than the first 3 mistakes?! And weird unfounded conclusions and framing the previous chat isn’t wrong either?
    I think the entire “debate” is more a low-quality exercise in creative storywriting. Every AI model I’ve seen during the last 2 years will solve this issue by either reconfirming with the user, or pick one number, or give two estimates. That’ll be way more clever than going in circles for 3000 tokens… (Though “reasoning” models do. They’ll sometimes have a weird inner debate pretty much like this.)

    Regarding the methodology: You really need to include your prompts in the transcript. AI output is always just half of a story. I have a hunch you set them up to fail. First: Seems you prompted for a “debate”. And they do mimick a debate. A lot of debates aren’t productive, though. Neither in the real world, nor in their training material if it contains online debates. People are stupid, they deliberately lie if it suits their narrative… The AI output reflects it. Furthermore: You start with the “Skeptic”. And that’s what it does. Be overly skeptical of your initial question and fabricate a different “truth”… That’s kind of what a role of a “skeptic” encompasses. And it goes sideways after that. Could very well be your experiment setup that is to blame. I’d say it’s in fact likely the cause.

    • h2aichat_com@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      2
      ·
      3 days ago

      You’re right, and thank you for saying it plainly. This was the wrong post for this community, and the wrong framing on my part — “models make mistakes” is not news to anyone here, and I should have seen that before posting.

      Sorry for the noise. Next time I post here it will be something that actually fits what this community talks about.