4. Tools close the rest
One
Retrieval takes one pass. An agent takes as many as the question needs.
It lists the parts, reads a property, traces an edge, and only then writes the answer. Each call is chosen after the last one came back, so the second question can be one the first answer raised.
Two
The paper's card reads 0.806 against 0.521, and both arms are the same loop.
Frozen card, the Apollo suite under Sonnet 4.5. Both arms run the same agentic loop; the winning arm was handed a report generator.
Nothing about the language model changed between those two numbers. What changed is the tool it was handed. The arm that could render a report scored where the other one did not.
Three
The second card, 0.900 against 0.569, is a different corpus and a different model.
Frozen card, the Eve corpus under Sonnet 4: a command-line tool against retrieval. This deck's own agent scores 0.90, Sonnet 5, the Apollo corpus, forty tasks.
It says the same thing about tools against one retrieval pass, and it is not the same measurement. Two cards, two labels, never averaged into one. The number this deck recorded for its own agent sits under them, labelled the same way.