Agent runtime comparator

Measure what the harness adds to the model.

Compare two agent setups across quality, latency, and cost to isolate the harness effect.

Try it now

See what the wrapper changes

Compare two agent setups with a practical weighting: task success first, then latency and run cost.

Harness A
Harness B

Live result

Harness B comes out ahead

A 11.4-point harness effect under this weighting.

Harness A88.4weighted score
Harness B99.8weighted score
Difference11.4points

One useful idea when the research moves. No noise.

Nine small tools for ideas just entering the field.