Agent runtime comparator
Measure what the harness adds to the model.
Compare two agent setups across quality, latency, and cost to isolate the harness effect.
Try it now
See what the wrapper changes
Compare two agent setups with a practical weighting: task success first, then latency and run cost.
Live result
Harness B comes out ahead
A 11.4-point harness effect under this weighting.
Harness A88.4weighted score
Harness B99.8weighted score
Difference11.4points
Frontier Tools