Editor's Note
Welcome to our inaugural issue. Many of you probably found us from the NYTimes article on tokenminnig. This is a list of resources, tips and tricks pulled together by the Neurometric.ai team - we build a token engineering platform that makes all of this easy. Thanks for reading.
This week’s resources
Choosing an Inference Engine in 2026 - link. This is relevant because choosing the best inference engine for your use case can lower inference costs significantly.
Cheaper Tokens Drive Higher AI Token Spend - link. A great case study in Jevon’s paradox and how lower token prices increase overall token spend.
Contenxt Engineering Guide 2026: Beyond Prompting - link. This article has some great tips on lowering token costs by compressing context, trimming prompt context that is no longer relevant, and caching.
SGLang 2026: High Performance LLM Inference - link. An overview of SGLang and how to optimize it.
MinMax Sparse Attention - link. Interesting research paper that could help inference costs.
VibeThinker 3B - link. A Small Language Model that beats frontier models on many coding tasks.
Case Study - AT&T
Venturebeat wrote an interesting case study about how AT&T used better SLM orchestration to lower their inference costs 90%. Read more about it here.
Thanks for reading this week. If you see thing that should be included, email me [email protected]
