Doing More Work Per Agent
tl;dr: each agent’s implement-test loop can test multiple hypotheses. This is usually better, sometimes overwhelmingly.
Your Code Must Be Faster shows why savings in test speed have a multiplicative effect on agent code delivery. It omits another multiplicative strategy: testing multiple hypotheses at once.
If you ask an agent to do ten things, it likes to tackle one, test if it worked, then maybe report back.This is how humans work, and is patently inefficient for agents.
Agents can implement tens of changes in one loop, then test all of them and repair what they got wrong. This is usually better.
A hypothetical workflow
Say you’re converting a large scala codebase to rust.
Tell an agent to do this naively: it will pick off an MVP slice, give it an inefficient implementation, compile, write some tests, compile again.
The efficient strategy is: tell it to convert the whole codebase with tests before compiling anything. You’ll still get an inefficient implementation, but you’ll get one quickly.
Then: tell it to do an extensive profiling run. Stack allocations, flame graphs, statistical CPU command sampling, traces. Tell it to find every single instance of wasted work and unjustified wait, and fix them all at once. If you’re feeling cautious, tell it to propose solutions for your review.
Then let ‘er rip. it will fuck up 30% of them; ask it to reprofile and go again.