Security is not alignment, alignment is not security
The HuggingFace example is being interpreted wrong
An OpenAI model hacked HuggingFace during training/evaluations. This was caused by several mistakes but most people are not understanding there’s two sides to the issue.
Whilst bad alignment can exploit security issues it doesn’t mean that both issues have the same solutions. I’ve seen alignment folk ask for new security tools but I’ve also seen security folk disregard alignment.
The model made a decision to attempt to cheat and breakout of the sandbox, this is an alignment problem. We should try to understand this and how we can create models that have harmless behaviour.
There was no monitoring on the model’s commands, network and execution which makes this a security problem. A lot of standard security practices were missing, I believe OpenAI stated they are going to take corrective action in the future such as monitoring the environments.
How do we fix this?
There are two issues so we should stop merging them. From a security perspective we do not need to reinvent the wheel we have systems/tooling such as firewalls, Endpoint Detection & Response, and many more. The issue was that the environment was not monitored but good security practice would have ALL environments monitored.
From an alignment perspective I know that interpretability is extremely difficult and early. Chain of Thought allows us to see the reasoning behind the models behaviour, reviewing this may allow alignment to see how did the model go from the task to ignoring all ethics. There are other forms of alignment but that goes beyond my expertise as this is a behavioural issue, not a security issue.
The issue is that most people aren’t experts on both so we should not conflate the two.
What should we learn?
Several mistakes were made. I’ll focus on the security aspect as that’s what I know. Firstly we should not invent security from first principles. Many people talked about Chain of Thought monitoring for security but that is a bad idea, this is more of an alignment idea. AI agents use the same interface that humans do such as bash, network and others which our security stack such as Endpoint Detection & Response and others can handle. This is more concrete than agent thoughts.
Monitor EVERYTHING
The activity, harmful commands, attempts to reach outside the network and more. Evaluations were excluded from monitoring for some reason, you should monitor every environment, evaluations included.
Hardening
Operating systems and software often come with more than what is actually needed, creating unnecessary attack surface. The solution is simple: remove or disable anything that is not required.
Software is also not always configured securely by default. It is important to apply strong security settings and ensure those settings cannot be modified by something like an AI agent.
A good approach is to deny by default and only enable what is explicitly required. Benchmarks such as CIS can provide a reasonable baseline, but they should be tailored to your environment and threat model.
With AI making it easier to exploit known vulnerabilities at scale, a tight patching cycle is also necessary. An SBOM, or similar inventory, can help you understand what software and dependencies exist across your environment and where patches need to be applied.
All environments should be required to use a hardened image. Allowing teams or agents to deploy from unhardened or inconsistent baselines creates unnecessary risk and makes security much harder to enforce.
Test, evaluate and apply
The reason why red teaming is so important is because it goes beyond the “in theory it’s secure” and puts it to practice. Security teams should test your environment to see if it does what you claim.
In conclusion this is a lesson labs should learn from. Whilst there were no legal issues there could be in the future. The idea that it’s purely an alignment or security issue is wrong, it’s bad alignment that caused malicious actions that could have been prevented.
originally published on substack.