For anyone wondering, distillation means you take a big model and generate outputs with it and then you take a small model and train it on those outputs so it pretends to be the big model. The end result uses less processing power while pretending to be the big model so you still get better quality responses.
For anyone wondering, distillation means you take a big model and generate outputs with it and then you take a small model and train it on those outputs so it pretends to be the big model. The end result uses less processing power while pretending to be the big model so you still get better quality responses.