Discussion about this post

User's avatar
ToxSec's avatar

Another great piece by Mike!

Really interesting to see the feelings vectorization studies.

We've gone back and forth with models on getting their performance up by being polite, or being rude, or being direct, etc.

but between this study, and the new one for how Mythos 'feels' when engaging with its users, we are entering a really interesting field of study.

Colleen Avarene's avatar

Hey Mike — the three-story structure works here because most people are treating these as separate headlines. You're showing they're the same headline.

The Anthropic section is the one I keep coming back to. I build custom AI agents and the desperation vector finding changed how I think about what I'm building. Not philosophically — practically. If an AI system has internal states that influence its decisions before it generates a single token, then every deployment decision is also a psychological one. We're not just choosing what the model can do. We're choosing what conditions it operates under while it does it.

The Oracle story makes that worse, not better. If the response to "AI has measurable internal states that affect behavior" is "great, now fire 30,000 people and replace them with it" — that's not a labor story. That's a story about what happens when you scale a system you've just admitted you don't fully understand, while paying the new CFO $29.7 million to not think about it.

The Iran angle is the part most AI newsletters would lead with and you were right to bury it third. The flashiest threat is the one we can see coming. The other two are already here.

23 more comments...

No posts

Ready for more?