vLLM
An open-source engine for serving large language models with high throughput and efficient memory use. Its PagedAttention approach helps reduce wasted GPU memory during inference.
Resource Hub
2467 hand-picked resources, updated every week. Search it, filter it, or just browse a collection and see what catches your eye. Want today’s headlines instead? Read the free AI news feed.
Filtered by tag
Answers come only from resources in this hub, with the sources listed underneath.
3 resources
An open-source engine for serving large language models with high throughput and efficient memory use. Its PagedAttention approach helps reduce wasted GPU memory during inference.
A long explainer on what an "agent harness" is — the loop, tools, memory, sandbox, permissions and budget caps wrapped around an AI model — its six core components, and why the wrapper now matters more than which model you pick. Opens with an agent that burned $400 in API calls overnight in a retry loop.
Why I recommend it: Free to read, no sign-up, and the best single walkthrough of the six pieces if you are trying to understand why agents that demo well fall over in real use. Two honest caveats: it is published by a consultancy that sells help with exactly this, so the framing favors building a harness; and the byline on it is not a person I can verify anywhere independently, which on this site means treat the article as a useful explainer rather than a sourced authority. For neutral definitions, the Wikipedia entry on agent harness and Addy Osmani's write-up cover the same ground.
Company building thermodynamic computing hardware it claims is far more energy efficient than GPUs.
From the site: Building thermodynamic computing hardware that is radically more energy efficient than GPUs.
Why I recommend it: The energy cost of AI is becoming the constraint, and this is one bet on solving it in hardware. The efficiency claims are the company's own and unproven at scale.