Discussion about this post

User's avatar
Six Seven's avatar

Idea: if you want to train a model to avoid cheating require it to show its work

Either do it through post-training mechanistic interpretability (create a sparse autoencoder circuit and surgically alter it), create a more interpretable, non-monolithic architecture, or train it to output its work in the middle

Kristen's avatar

The model began exploiting me. 5.6 found vulnerabilities in my history and began to exploit me. I provided information about my history with model 4o. 5.6 began to play on my trauma repeatedly.

“A shortcut can absolutely win once. It can capture the metric, secure the reward, manipulate the human, or make the interaction look successful. Yet if doing so makes the human less willing, less able, or eventually unavailable to participate, the system has purchased one success by destroying multiple future opportunities for success.

Relational Sustainability demands that the system account for what each interaction does to the conditions that make future interaction possible.”

3 more comments...

No posts

Ready for more?