中文
Agent infrastructure

beellama.cpp

anbeeld/beellama.cpp

llama.cpp fork focused on local LLM inference efficiency, using KVarN, KV-cache precision tailoring, and low-bit quantization to support longer context with better precision in the same VRAM.

Topics

  • ggml
  • kv-cache
  • llama-cpp
  • llm-inference
  • quantization
  • inference
  • llm
  • llm-serving