Show HN: Find the best local LLM for your hardware, ranked by benchmarks | Hacker News
TL;DR AI
2 min readKey summary
A Show HN post introduced a tool that ranks local LLMs by benchmark performance to help users pick models for their hardware.
Commenters said the ranking should also surface quantization quality loss, maximum context, and token-generation speed under long contexts.
They also wanted batch-parallelism effects, KV cache quantization, and other implementation details that affect real-world VRAM and throughput.
The discussion reflected a broader need for hardware-specific guidance, especially for Apple Silicon and fast builds like MLX.



