Discussion about this post

User's avatar
raghav's avatar

>> it could enable the model to distinguish between its training and deployment stages, meaning it could know to behave in a way that humans will approve of in the former, but act very differently in the latter.

How is this possible? Isn't the behavior shown at the training rewarded/reinforced, and hence the same behavior will be displayed in deployment. I understand that situational awareness can make a model behave differently during evaluation. But during training, the model becomes what is reinforced right?

Shubbair's avatar

amazing details are there.

No posts

Ready for more?