Batching by Length Instead of Looping Item by Item for SLM Optimization
Previous articles in this series discussed constraining output space as well as reusing the prompt prefix with a key-value cache, both framed as approaches to small language model (SLM)
Read More
