Help needed: Native 64k+ context GGUF model for llama.cpp

Inner-End7733 · reddit · 2026-07-31

A developer encountered a context bottleneck while trying to fully localize Hermes Agent on low-spec hardware (64GB RAM + RTX 3060).

The Problem:

The Request:

Looking for a dense, non-reasoning model natively supporting 64k+ context, already quantized in GGUF, to meet the strict requirements of the local agent workflow.

Original post →

More from coding & agent

coding & agent channel →