Side by side
Compare two AI models
Pick two models and see what each costs, how well it works and what it's actually for. The scores are my own judgment on a 1–5 scale from using these — not vendor claims and not benchmark results. Your picks stay in the address bar, so you can send the comparison to someone else.
Choose a second model above to see them side by side.
Chat and everyday writing
Hermes Agent
Hermes
VisitFree value
Performance
Range of uses
Free and open, rougher than the commercial tools — good for learning agents.
- Published benchmark scores
- MMLU-Pro82.9%
hard multiple-choice questions across professional subjects · tested: Hermes 4 405B
Benchmark List, citing the Hermes 4 report · 13 August 2025 report · maker's own figures
- GPQA Diamond72.7%
graduate-level science questions · tested: Hermes 4 405B
Nous Research technical report · 13 August 2025 · maker's own figures
- Artificial Analysis Intelligence Index9
combined score across ten reasoning and knowledge tests · tested: Hermes 4 405B
Artificial Analysis · checked September 2026 · independent
Far below the hosted flagships — this is the price of running it yourself.
The agent itself has never been benchmarked. These are scores for Hermes 4 405B, the model it runs on, and the headline numbers come from Nous Research's own technical report.
Figures copied from the sources above, checked 17 September 2026.
- Cost
- Free, MIT-licensed and self-hosted; you pay your own model provider.
- Performance
- An open agent you point at a task rather than a chatbot; capable but rougher around the edges than the commercial assistants.
- Use it for
- Automating a repeated multi-step task
- Learning how agents actually work
