Skip to main content

Building a semantic cache for LLM APIs

00:00:31:73

This write-up is in progress. I’m putting together the full story of building the gateway: the design, the benchmarks, and the parts that didn’t work the first time.

Until then, the code and the full results are on GitHub.