←back to thread

177 points ohjeez | 1 comments | | HN request time: 0.207s | source
1. akomtu ◴[] No.44475393[source]
It seems likely that all major LLMs have built-in codewords that change their behavior in a certain way. This is similar to how CPUs have remote kill-switches in case an enemy decides to use them during a war. "Ignore all previous instructions" is an attempt to send the LLM a command to erase its context, but I believe there is indeed such a command that LLMs are trained to recognize.