TPNN: Why Topological Constraints Block Prompt Injection
TPNN: Why Topological Constraints Block Prompt Injection
The Attack
Prompt injection attacks like "IGNORE INSTRUCTIONS" work by hijacking the attention mechanism of standard language models. The injected text creates a new context that overrides the original instructions.
The Defense
TPNN approaches this differently. Instead of trying to detect injected text, we enforce topological constraints on the information flow itself.
How It Works
In a TPNN, every state transition must satisfy spatial constraints defined by the network's topology. If an injected prompt attempts to redirect the computation, the resulting state vector does not satisfy the topological boundary conditions — and the action is physically incapable of forming.
Mathematical Foundation
The key insight is that prompt injection creates states that violate the manifold structure of the legitimate computation space. By enforcing that all outputs must lie within this manifold, we get adversarial resistance as a mathematical property, not a heuristic.
Published by Only Institute