Topic
Evaluation
4 stories on this topic, newest first.
ModelsExplainer
How to read a model card without getting fooled
Model cards are part specification, part marketing. Here is a field-by-field guide to the numbers that matter, the ones that are routinely gamed, and the questions a card should answer before you ship on it.
ResearchExplainer
What "reasoning" models actually do differently
Reasoning models are trained to spend tokens thinking before they answer. Here is what that training involves, why it works on some problems and not others, and how to decide when to pay for it.
IndustryEditorial
Benchmarks became a marketing channel. Treat them like one
Benchmark tables were meant to be measurements. They are now launch collateral, optimised for by every lab and reported under whatever settings look best. That does not make them useless, but it changes how they should be read.
IndustryExplainer
The real cost of running an AI product, line by line
Token spend is the line everyone watches and rarely the largest. A working breakdown of where the money goes in a production generative AI product, from inference and evaluation to the humans in the loop.



