Topic

Evaluation

4 stories on this topic, newest first.

  1. ModelsExplainer

    How to read a model card without getting fooled

    Model cards are part specification, part marketing. Here is a field-by-field guide to the numbers that matter, the ones that are routinely gamed, and the questions a card should answer before you ship on it.

    4 min read

  2. ResearchExplainer

    What "reasoning" models actually do differently

    Reasoning models are trained to spend tokens thinking before they answer. Here is what that training involves, why it works on some problems and not others, and how to decide when to pay for it.

    4 min read

  3. IndustryEditorial

    Benchmarks became a marketing channel. Treat them like one

    Benchmark tables were meant to be measurements. They are now launch collateral, optimised for by every lab and reported under whatever settings look best. That does not make them useless, but it changes how they should be read.

    3 min read

  4. IndustryExplainer

    The real cost of running an AI product, line by line

    Token spend is the line everyone watches and rarely the largest. A working breakdown of where the money goes in a production generative AI product, from inference and evaluation to the humans in the loop.

    3 min read