Topic

Inference

3 stories on this topic, newest first.

  1. NewsAnalysis

    Inference prices keep falling. Here is who actually benefits

    Per-token prices for frontier-class models have dropped by roughly an order of magnitude every 18 months. The savings are real, but they land unevenly across the stack.

    4 min read

  2. ResearchAnalysis

    Test-time compute changed the scaling roadmap. Here is what it costs

    For a decade, progress meant bigger training runs. Now labs can trade inference compute for capability instead. That shifts the economics from capex at the lab to opex at the user, and it changes what "a better model" means.

    3 min read

  3. IndustryExplainer

    The real cost of running an AI product, line by line

    Token spend is the line everyone watches and rarely the largest. A working breakdown of where the money goes in a production generative AI product, from inference and evaluation to the humans in the loop.

    3 min read