Default avatar
npub1uqee...jckg
npub1uqee...jckg
Microsoft wrote a paper about how much LLMs degrade in multi-turn convos. I’m glad my hunch was finally confirmed. I’ve felt for a while that resetting the context and tweaking the first prompt is often far more effective than clarifying with a follow up prompt when the LLM misunderstands or fails to solve the problem on the first attempt. My guess was that there just aren’t enough examples of multi turn convos in the RLHF training dataset, especially of LLMs recovering from a mistake, but I hadn’t considered how benchmarks play into it. Because the format of most benchmarks is just “question->answer” without any follow ups your LLM can test well and still fall apart the moment the user asks a second question.
Google’s latest version of Gemini 2.5 Pro is freakily smart. I was going to send an email pointing out where I thought someone was incorrect about the ability of moral facts to affect our brains backed up with research. I fed the email through Gemini 2.5 Pro to review it and after several back and forths became convinced that I was actually incorrect and completely misunderstanding the core of the argument.
Even Anthropic’s experts who work on their AI models are filing AI generated court documents with hallucinations in their own court case. This happens all the time but normally it’s done by lawyers who don’t really understand the technology and don’t even know it can hallucinate. Here it’s inexcusable because this guy makes these models and should really know better.
Random thought, why do we have a sex offender registry for only sex related crimes? I suppose I would want to know if I guy who molested kids moved in next door but I would also want to know if a guy who broke into multiple homes moved in next door. I don’t see what distinguishes sex related crimes from other crimes that warrants only sex crimes getting a registry. We should pick a standard and apply it equally, either every criminal goes on a public registry or no one does.
So after everyone (rightfully) freaking out about Kamala Harris imposing price controls on drugs now Trump has done the exact same thing and instituted drug price controls? I’m giving up. Nothing in the timeline we’ve fallen into makes sense anymore.
ChatGPT taking every opportunity to trip over itself complimenting how right and intelligent you are does have its upsides. Model refusal rates seem to be significantly lower.
Wow this is really really bad. How did this even pass the giggle test? Of course backing up encrypted messages to a plaintext file is stupid. What’s even the point of encrypting them at all at that point?
I’m sitting next to an old man in a wheelchair who’s spent the last half hour watching the most mindless TikTok slop with the volume at full blast. There are so many things I don’t understand. I would be embarrassed to play anything out loud on my phone in a public place and would be mortified if it was brain rot. Apparently he just doesn’t care.
I used an LLM to brainstorm some deviously persuasive rhetorical tactics for an email and was pleased until I realized that they'll probably be used on me soon. The average human is not ready for intelligence equivalent to a top performing human to be unleashed on them to tailor every message and request to their precise idiosyncrasies.