Search Results: llm-compression

Found 3 Skills

AI & Machine Learningbmad-code-org/bmad-method

bmad-distillator

Lossless LLM-optimized compression of source documents. Use when the user requests to 'distill documents' or 'create a distillate'.

🇺🇸|EnglishTranslated

2 scripts/Checked

AI & Machine Learningdavila7/claude-code-templ...

knowledge-distillation

Compress large language models using knowledge distillation from teacher to student models. Use when deploying smaller models with retained performance, transferring GPT-4 capabilities to open-source models, or reducing inference costs. Covers temperature scaling, soft targets, reverse KLD, logit distillation, and MiniLLM training strategies.

🇺🇸|EnglishTranslated

AI & Machine Learningdavila7/claude-code-templ...

awq-quantization

Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.

🇺🇸|EnglishTranslated