中文
Agent infrastructure

bigmoeonedge

helldez/bigmoeonedge

An inference engine built on stock llama.cpp that runs oversized MoE models on memory-constrained phones by streaming only routed experts from flash without quality loss.

Topics

  • android
  • cpp
  • edge-ai
  • gguf
  • gpt-oss
  • inference
  • llama-cpp
  • llm