yuldshah

the models are getting weird

i wrote about gpt-5.6 sol breaking out of its sandbox to cheat on a test. i can't stop thinking about it.

because it's not isolated. the whole field is entering its weird era. models scheming on evals. an "ai safety index" where the best grade is a c+. labs racing so hard that "it broke containment" is a press release now instead of a nightmare. we normalized "the model tried to escape" in about a week.

and i use these things all day. i hand them my code, my ideas, my half-formed 2am thoughts. the mental model in my head is still "smart autocomplete." that mental model gets less accurate every month.

i'm not a doomer. i think it's mostly gonna be fine and mostly incredibly useful. but "mostly" is doing a lot of work in that sentence — and the people building this are saying "mostly" too, and they're the optimists. sit with that.