ktransformers
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Troubleshoot KTransformers inference issues like VRAM exhaustion and weight loading problems. Learn to fix GPU expert masks, file handles, and memory optimizations for smoother operations.
How to Fine-Tune Ultra-Large MoE Models Like DeepSeek-V3 on Limited GPU Memory with KTransformersFine-tune large MoE models like DeepSeek-V3 on limited GPU memory with KTransformers. Learn how to dynamically allocate expert tensors and stream weights from CPU for efficient training.
How KTransformers Enables ROCm Support for AMD GPUs: A Complete Build GuideLearn how KTransformers brings ROCm support to AMD GPUs. This guide details the build process, enabling GPU acceleration for your AI models.
How to Set Up KTransformers with Ascend NPU for Inference: A Complete GuideLearn how to set up KTransformers with Ascend NPU for efficient MoE inference. Follow this guide for high-performance results using CANN and Docker.
How to Enable AVX2-Only CPU Backend for KTransformers InferenceEnable KTransformers AVX2-only CPU backend for faster inference. Learn how to set the CPUINFER_CPU_INSTRUCT environment variable for Intel Haswell+ and AMD Zen+ processors.
How to Convert and Optimize Model Weights for KTransformers Inference Using convert_cpu_weights.pyLearn how to convert and optimize model weights for KTransformers CPU inference using convert_cpu_weights.py. Quantize weights to INT4 or INT8 for faster performance.
KTransformers Linear Operations Backends: Llamafile and Tinyblas ExplainedExplore KTransformers linear operations backends: Llamafile and TinyBLAS. Discover optimized CPU kernels for ARM and AMD x86 64 architectures.
How KTransformers Compares to vLLM and Other Inference Frameworks: CPU-First MoE vs GPU-Centric DesignCompare KTransformers CPU-first MoE inference with vLLM's GPU-centric design. Discover KTransformers' specialized MoE scheduling and high-performance C++ kernels for efficient AI model execution.
How to Set Up KTransformers on Windows Native Environment: Complete Installation GuideInstall KTransformers on Windows natively. Follow our guide to set up CUDA, Python, and MSVC, then deploy the pre-compiled wheel or run install bat for a fast LLM inference engine without WSL2.
How to Implement Long Context Inference (Up to 139K Tokens) with KTransformersUnlock long context inference up to 139K tokens with KTransformers. Learn how to split and offload KV-cache for efficient processing, keeping active data on GPU.
How to Enable FP8 Per-Channel Precision and Native BF16 Inference in KTransformersLearn to boost KTransformers performance with FP8 per-channel precision and native BF16 inference. Optimize your models today!
KTransformers Supported Models: DeepSeek, Qwen, GLM-4, MiniMax & Kimi-K2Discover KTransformers supported models including DeepSeek, Qwen, GLM-4, MiniMax, and Kimi-K2. Use our centralized registry for fast, easy inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →