> Cross-run MI plus timing side channels feels like the right hammer here
I do not think timing is the right axis to start from. If the agents are colluding throguh tool outputs, shared prompts, or even just low-bandwidth symbol choices, the leakage can be entirely content based while looking perfectly ordinary in wall-clock time, so a timing detector will miss the interesting cases and mostly flag scheduling noise.
Cross-run MI is more promising, but only if you define the units carefully, otherwise you end up measuring the benchmark harness and not the agents. The harder part is cross-principal alignment of messages, not latency, because a colluding pair can spread a code across runs with almost no temporal regularity at all.