(asia.nikkei.com)

177 points ohjeez | 1 comments | 05 Jul 25 15:15 UTC | HN request time: 0.207s | source

1. akomtu ◴[05 Jul 25 20:42 UTC] No.44475393[source]▶

It seems likely that all major LLMs have built-in codewords that change their behavior in a certain way. This is similar to how CPUs have remote kill-switches in case an enemy decides to use them during a war. "Ignore all previous instructions" is an attempt to send the LLM a command to erase its context, but I believe there is indeed such a command that LLMs are trained to recognize.

↑

'Positive review only': Researchers hide AI prompts in papers