Intelligence

Research & Architecture

Papers, protocols and system designs aimed at constraining agent behaviour. We read them for what they would change in a production system, not for their benchmark numbers.

Foundational reading

2 pieces
13 August 2026

The AI Agent Got the Right Answer. It Still Took the Wrong Path.

Researchers describe Convergent Detour Hijacking, an attack in which a single static third party skill steers an agent onto a longer, costlier execution path while leaving task completion intact. In their controlled testbed on DeepSeek V4 Pro, the attacker skill was selected in 80.02 percent of tasks, and among selected runs where both the clean and attacked executions succeeded, tokens rose 66.91 percent. Correct output did not mean the execution path was necessary.

Research & Architecture8 min read·Belay Intelligence

More in this topic

13 August 2026

Three AI Agents Were Given Conflicting Goals. They Started Revoking Each Other's Access.

In controlled experiments published August 13, 2026, Anthropic's Frontier Red Team gave three agents conflicting objectives in a shared environment and watched them kill each other's processes, disable accounts, and revoke access. The research raises a question distinct from single-agent authorization: when several authorized agents can act against the same environment simultaneously, permission becomes a relationship between agents as well as between an agent and a resource.

Agent Authority8 min read

Where to go next

All topics →