Skip to main content

Dec 16, 2025

Prompt caching: 10x cheaper LLM tokens, but how?

6,066 words

  • Quantization from the ground up

    A complete guide to what quantization is, how it works, and how it's used to compress large language models

  • The new ngrok.ai

    The ngrok AI Gateway is now app.ngrok.ai. One key, one URL to access any model, including the ones you run yourself, with access controls and per-call cost visibility.