ai-agents
You're Arguing About the Wrong Layer
Researchers ran three models through three harnesses on 100 coding tasks. The harness moved performance almost eight times more than the model did, and the model rankings flipped six times out of nine. Your model-choice argument is measuring the scaffolding.