"... Using tree-of-thought reasoning allows TAP to navigate a large search space of prompts and pruning reduces the total number of queries sent to the target. In empirical evaluations, we observe that TAP generates prompts that jailbreak state-of-the-art LLMs (including GPT4 and GPT4-Turbo) for more than 80% of the prompts using only a small number of queries..."
I watched the latest Karpathy video, where he shows that you can create a prompt suffix that looks like jibberish, using an iterative approach. It looks like this technique will be impossible to defend against, because it discovers for an arbitrary string that causes the jailbreak.
Looks like this paper is building on the token level attack from that work.