I’m giving a talk tmr at PhilML workshop at 10:35am! Please note that they’ll be actual philosophers (other speakers!) at the workshop too, so come by for some spicy thinking questions for you to ponder as you travel back home... :) It’s a great way to end icml.
Flying to ICML ✈️ I’ll be spending the week either at the conference or at restaurants/street food stations eating at least 4 meals a day. Also giving a keynote at PhilML workshop on 11th:)
Prompt engineering is still a black box. Why does changing X drastically change Y? Are there governing rules behind this evolution? Our new work proposes a simple way to uncover factors that might matter when refining prompts 👇
Thrilled to share that our paper on "Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits" has been accepted at AISTATS 2026! 🚀🚀
Read more about how input mutations can be mapped to interpretable behavioral insights.
arxiv.org/abs/2602.00092
🧵
I got my account back! Thank you, first and foremost, to everyone—friends, GDM colleagues---who personally alerted me to this incident and retweeted that I'm hacked, as well as folks at X who helped me regain access. While this incident was terrible (I heard the scammers made
Tomorrow 9:30am #NeurIPS2025 Room 30A-E I'll talk about " 📈Towards Pareto frontier of interpretability:
15 years of interpretability research in 15 mins"🚅
@ mech interp workshop mechinterpworkshop.com