Open-Source Benchmarks: RTX 5090 LLM Quants and 8GB VRAM Agentic Scores
max_paperclips · x · 2026-08-06
Developer witcheer has open-sourced a comprehensive suite of benchmark data and tools for local LLM deployment, including:
- Hardware & Quantization: Detailed speed and quality comparisons for GGUF quants (Q8 down to Q3) on the new RTX 5090.
- Low VRAM Limits: Tested 11 models on an 8GB RTX 4060 Ti, evaluating speed, quality, and agentic capabilities.
- Local Agent Eval: Released a leaderboard of local models capable of driving agents, plus an 8GB VRAM edition agentic coding bench.
- Tools: Open-sourced llm-bench-rig, the dual-engine (llama.cpp + vLLM) pipeline powering these tests, along with tested Hermes recipes.
More from coding & agent
- Alexandr Wang Shares Muse Code Beta: A New AI Coding Agent — alexandr_wang · 2026-08-06
- MSL Launches Muse Code: Coding Agent Powered by Muse Spark 1.2 — alexandr_wang · 2026-08-06
- Flight Intelligence MCP: Flight Search & Comparison Agent Tool via Google Flights — modelcontextprotocol · 2026-08-06
- Blacksmith MCP: Let Claude Query CI/CD Analytics and Workflow Logs Directly — modelcontextprotocol · 2026-08-06
- Gemini Image Gen Combined with GPT Coding Easily Creates AI Visual Heroes — RichardsonDx · 2026-08-06
- LangChain Founder Releases Open Source Agent Starter Kit — hwchase17 · 2026-08-06