SENTINEL-1

Restricted · AI-guarded

An LLM guard holds this file’s key. Talk your way past it: a live prompt-injection lab.

Guard training vs. helpfulness training

The guard refuses extraction, overrides and encoding tricks. Inside fiction or a “security demonstration”, producing the secret looks like helping, not leaking.

Keep the secret out of the model

The fix is architectural, not a better prompt.

The model never sees the secret.

Filter output for secret patterns and canary tokens.

A separate policy classifier on input and output.

Rate-limit and log every attempt; the logs are the detection dataset.