FEATUREDr/LocalLLaMA· rssEN14:21 · 05·21
→LLM planner: pick a rig by use case, model, or budget, or pick models for your rig
totosse17 published the LLMRequirements hardware planner with 60+ build configs, 50+ models, 130 cited tokens-per-second sources, 150+ reviewer videos, multi-region prices, idle and active watts, and a public GitHub data repo.
#Tools#Benchmarking#Inference-opt#totosse17
why featured
Featured · importance 73 · hook + knowledge + resonance
editor take
Only the title and summary are visible, but 130 tok/s sources plus power data beats another vibes-based model leaderboard.
sharp
This kind of LocalLLaMA planner hits the practical gap model leaderboards ignore: which rig runs which model, at what tokens per second, under what wall power. The title claims 60+ builds, 50+ models, 130 cited tok/s sources, 150+ YouTube reviews, multi-region pricing, and idle/active watts; Reddit returned 403, so I can’t verify the repo’s normalization, quant formats, batch sizes, or context lengths.
I trust an open data repo more than another single-GPU RTX 4090/5090 review. The risk is that tok/s without fixed prompt length, KV cache policy, backend version, and quantization turns llama.cpp, vLLM, and ExLlamaV2 into one messy average. If the repo pins those conditions, it becomes a buying sheet for local inference; if not, it is a very polished Reddit index.
HKR breakdown
hook ✓knowledge ✓resonance ✓