Lesson 4.213mIntermediate12.8k students

Managing cost and latency

Cost is prompt design, model choice, and how often you call. Measure per-request tokens before you try to optimise anything.

This lesson sits in Shipping to Production, part of Building AI Apps with LLMs. It assumes what came before it and leads directly into the next lesson in the module.

In this lesson you will

  • Attribute spend to specific calls and prompts
  • Trim prompts and route by difficulty
  • Set timeouts and fall back gracefully

Resources