中文
Agent infrastructure

lucebox

luce-org/lucebox

Fast LLM inference server for consumer-grade GPUs that uses speculative decoding and custom kernels to raise throughput while lowering memory use.

Topics

  • kernel
  • llama-cpp
  • local-ai
  • qwen
  • rtx3090
  • megakernel
  • cuda
  • cuda-kernels