Our LLM bill tripled in 6 months, i need to lower it

We run a B2B support bot plus some internal agents. Token usage went up, which is good but the bill went up faster than revenue, which is bad=)

Things I already know about but haven't done properly:

- Sending simple intents to smaller models and keeping frontier models for the hard ones

- Caching repeated prompts (we have a lot of near-identical system prompts)

- Committing to volume in exchange for a discount

I'm looking at openrouter of course but also at llmapi. ai because it puts routing, caching and usage limits all together or something similar, and seen the bill actually drop?

I'd love real numbers if you're willing to share them even rough ones

Author: Repulsive-Koala6511