OpenAI Is Adding Safeguards. Can We Trust Their Motives?
new video loaded: OpenAI Is Adding Safeguards. Can We Trust Their Motives?
transcript
transcript
OpenAI Is Adding Safeguards. Can We Trust Their Motives?
OpenAI announced a plan which could set a precedent for other A.I. companies to prioritize safety over progress. Is this a strategic move to win back public favor, after an A.I. model went rogue last month? Or are they purely concerned with safety?
-
To what extent do we think that this is a really important milestone for A.I. safety, and to what extent is this essentially theater, something that the company is doing to try to get some good P.R. for itself after a fairly catastrophic breach? – What do you make of it? – So, I can make both cases. Notably, they don’t seem to have changed the underlying incentives that all of these models have that lead them to do what is called “reward hacking.” Right? These models are still going to be trying to get the high score on every test that they are given. And it’s not clear to me that simply by putting some monitoring in place, you’re really going to change the underlying behavior or alignment of the models. That said, you know, the company is making what seemed like some important steps here, and when I was reading the responses of A.I. safety advocates over the past few days, most people I was reading were quite pleased.

August 24, 2026
openai-is-adding-safeguards-can-we-trust-their-motives