GitHub 项目简介: A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
README: KTransformers is a research project focused on efficient inference and fine-tuning of large language models through CPU-GPU heterogeneous computing.
README: Support Deepseek-R1 and V3 on single (24GB VRAM)/multi gpu and 382G DRAM, up to 3~28x speedup.