Arena Blog (LLM and Agent Evaluation Research)
Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.
Why I recommend it: Free to read. The clearest writing anywhere on why two people using the same model get very different results — the tool wrapped around the model changes the cost and the outcome. They run the leaderboards they write about, so treat their rankings as one measurement, not the verdict.
