An Empirical 8-Model Benchmark on Agent Harness Architecture, Model Over-Refusal, Token Efficiency, and Enterprise Sovereignty in Modern AI-Driven Penetration Testing 1. The Core Thesis: Offensive Capability Is a Function of the Harness, Not the Raw Model In my book...


