Skip to content

vllm.models.kimi_k3.nvidia

Modules:

  • dspark_mla

    K3 dense MLA draft model for DSpark speculative decoding.

  • kda
  • kda_metadata

    Kimi-K3 specialization of GDN attention metadata.

  • latent_moe_runner
  • low_latency_gemm

    Kimi-K3 decode GEMM selection for unquantized BF16 on SM103.

  • mla

    Clean Multi-head Latent Attention for Kimi-K3 (NVIDIA).

  • model

    Kimi-K3 multimodal model implementation for vLLM.

  • mtp

    Inference-only Kimi-K3 Multi-Token-Prediction (MTP) draft model.

  • ops