NewsAnalysis
Inference prices keep falling. Here is who actually benefits
Per-token prices for frontier-class models have dropped by roughly an order of magnitude every 18 months. The savings are real, but they land unevenly across the stack.
Latest edition·
The stories, launches and research that actually matter. Reported fast, explained properly, with the noise cut.
NewsAnalysis
Per-token prices for frontier-class models have dropped by roughly an order of magnitude every 18 months. The savings are real, but they land unevenly across the stack.
Latest
Archive →Section
2 stories
Frontier and open-weight model releases, benchmarks, pricing and what each one changes for the people who build on them.
Section
2 stories
Papers, techniques and results from labs and universities, translated into what they mean in practice.
Section
2 stories
Developer tooling, agents, evaluation, gateways and the software stack around generative AI, tested and compared.
Section
2 stories
The economics, strategy and regulation of generative AI: who is spending, who is winning and why.
NewsExplainer
General-purpose AI model obligations under the EU AI Act have applied since August 2025. A year later, here is what providers have actually had to do, and what still applies to you if you only fine-tune or deploy.
ModelsExplainer
Model cards are part specification, part marketing. Here is a field-by-field guide to the numbers that matter, the ones that are routinely gamed, and the questions a card should answer before you ship on it.
ResearchExplainer
Reasoning models are trained to spend tokens thinking before they answer. Here is what that training involves, why it works on some problems and not others, and how to decide when to pay for it.
ResearchAnalysis
For a decade, progress meant bigger training runs. Now labs can trade inference compute for capability instead. That shifts the economics from capex at the lab to opex at the user, and it changes what "a better model" means.
ToolsComparison
LLM gateways sit between your applications and model providers to handle routing, keys, failover, budgets and logging. We compare Bifrost, LiteLLM, Portkey, Kong AI Gateway and Cloudflare AI Gateway on the decisions that actually differ.
ToolsSurvey
Agent frameworks have split into graph-based orchestrators, lab-native SDKs, multi-agent role systems and typed minimalists. This survey maps the landscape, the design bets behind each camp and the questions to ask before committing.
IndustryEditorial
Benchmark tables were meant to be measurements. They are now launch collateral, optimised for by every lab and reported under whatever settings look best. That does not make them useless, but it changes how they should be read.
IndustryExplainer
Token spend is the line everyone watches and rarely the largest. A working breakdown of where the money goes in a production generative AI product, from inference and evaluation to the humans in the loop.