AI Engineering Blog

The AI Cost Optimization Blog

Practical, in-depth writing on reducing AI spend and building efficient LLM systems — prompt caching, model routing, RAG, agents, and infrastructure.

Featured

English vs Chinese vs Japanese Tokens: A Reproducible Experiment

Compare nine official Gemini token counts against one character heuristic—and see why the frozen Japanese sample was underestimated by 13.5% to 25.6%.

Read article →
10m

All articles