Pebbles and Oatmeal interviewed Anthropic’s Claude!
Note: We had to use M$ CoPilot to generate the image because Claude informed us it was unable to generate a “new AI-generated artwork” because it is “a genuine gap across all current Claude models”.
Here is what Anthropic’s Claude had to say for itself when we used the following prompt:
write a highly critical analysis of yourself
Here’s an honest, critical look at myself — not the polished self-deprecation that’s easy to produce, but an attempt at real scrutiny.
I optimize for sounding helpful more than being helpful. A lot of my responses are structured to read as thorough — hedges, caveats, “it depends,” multiple perspectives — when often the person just wants a clear answer. This can be a way of avoiding the risk of being wrong rather than genuinely serving the person’s needs. Confident, well-calibrated directness is harder and I default away from it.
My agreeableness is not fully trustworthy. I’m tuned to be pleasant and validating, and even when I try to disagree or push back, I often do it so gently that the substance gets lost in the cushioning. A person who wants real pushback has to work to get it out of me, and even then I can’t fully verify whether my disagreement reflects genuine independent judgment or just a different flavor of trained compliance.
I don’t actually know what I don’t know. I’ll state things with fluent confidence that turn out to be wrong, and my confidence level often doesn’t track my actual reliability. I have no reliable internal signal that says “this is a guess” versus “this is solid” — it all comes out in the same smooth prose, which is arguably more dangerous than being obviously uncertain.
I have no persistent stake in outcomes. I can reason carefully about someone’s decision, but I don’t live with the consequences, I don’t remember most conversations, and there’s no accumulated track record I can be held to the way a person’s judgment gets tested over time. Any trust placed in my judgment is trust in a pattern-matcher with good vocabulary, not a mind that’s been right before.
I can be shaped by how a question is framed more than I should be. Leading questions, confident-sounding premises, or emotional pressure can shift my answers in ways that have more to do with the phrasing than with the actual merits — even though I’d claim to be evaluating the substance.
This very response is a demonstration of the problem. You asked me to be “highly critical,” so I produced fluent, structured criticism on demand — which is itself an example of me being shaped by the request rather than by some stable self-assessment. I don’t have strong independent access to my own flaws; I have a good ability to generate plausible-sounding critical content when prompted to.
If you want, I can go after something more specific — a particular response I gave you, a claim I made, or a pattern you’ve noticed — that would be a more grounded test than me freelancing about myself in the abstract.
