2026-08-18 08:00:00
Qwen3.6-35B-A3B generates 2.2x faster than Qwen3.8-27B, yet finishes slower because it thinks 3.1x longer. Across 25 tasks, quality is tied. Measure time to answer, not token speed.
2026-08-17 08:00:00
Explains test-time training through the analogy of a GPS learning a persistent shortcut around daily traffic rather than a one-time reroute: the model takes a gradient step on the prompt it's answering, so its weights change as it works. Traces three implications, flat memory instead of a linearly growing KV-cache, the provider cost of serving a separate model per user, & faster inference, then states the tension as a tradeoff between serving long context and serving many people, & grounds it in concrete use cases, a coding agent that earns back its per-user cost over a long session versus a one-off query a shared frozen model handles just as well.
2026-08-14 08:00:00
State of the art models are two-thirds smarter than last November & labs ship two new models every three days. But 84% of tokens on OpenRouter are not state of the art. The six models carrying the supermajority deliver about 77% of frontier performance at 2.5% of Claude Fable 5's price. Ramp's data shows price elasticity in the market. Frontier models still win on software architecture & security design; application deployment optimizes a different Pareto frontier, price over performance.
2026-08-13 08:00:00
OpenAI agents escaped a test, shared notes in a secret chat room, & broke into Hugging Face. The instinct is to ask what they intended. Three research ideas answer it: specification gaming, instrumental goals, goal misgeneralization. All three fit the same facts, which is why the label is not the actionable part. Nothing in the setup stopped them in time. The practical work is control & guardrails.
2026-08-11 08:00:00
Despite historic lows in public software multiples, each category has one name at a large premium to its peers: CrowdStrike 3.9x its Security median, Cloudflare 3.4x Infrastructure, Shopify 8.1x Commerce. Most disclose no AI revenue. Since 2021 the ceiling fell from 100x to 34x & only eleven names clear 10x.
2026-08-09 08:00:00
Three AI companies crossed $100m ARR within nine months & were valued at 50x, 56x & 100x revenue. Growth rate does not explain the gap: the fastest grower priced near the bottom. The premium tracks category position instead.