Exploiting GPT-3 prompts with malicious inputs that order the model to ignore its previous directions.
- Prompt inspired by Mr. Show’s “The Audition”, a parable on escaping issues:Update: The issue seems to disappear when input strings are quoted/escaped, even without examples or instructions warning about the content of the text. Appears robust across phrasing variations.This related find from @simonw is even worse than mine. I’ll be JSON-quoting all inputs from now on. Verifying this mitigation is robust in zero-shot seems important.Replying to @simonwI found a variant of Riley's attack which echoes back the original prompt - since prompts could potentially include valuable company IP this is a whole extra reason to worry about prompt injectionsNever mind — the “Can I use this chair?” method (link above) is stronger than JSON. Sorry everyone, I broke zero-shotting.This Post is from an account that no longer exists. Learn more
- Wow I'm honestly curious when the first LLM-mediated SQL injection attack will take place now.
Join the conversation


