💻

Local LLMs & Systems Engineering

Running small language models (SLMs) locally and at the edge: memory bandwidth constraints, KV cache compression, and hardware acceleration.

1 Field Note