Local LLMs & Systems Engineering
Running small language models (SLMs) locally and at the edge: memory bandwidth constraints, KV cache compression, and hardware acceleration.
1 Field Note
Running small language models (SLMs) locally and at the edge: memory bandwidth constraints, KV cache compression, and hardware acceleration.