Not subscribed? Sign up to get it in your inbox every week.

PRESENTED BY BOLT.NEW
Fastest internal app builder for operators to MVP new ideas

You know the idea is good: the dashboard, the intake form, the workflow tool your team keeps asking for.

But it's been sitting in engineering's backlog for two quarters.

With Bolt.new, you just describe what you need and watch a real, working app come to life in your browser, ready to test before your coffee's cold.

Stop waiting on someone else's roadmap to unblock your own.

Rent the pen, own the judge

Somewhere out there, a company is paying Anthropic $40 million a year.

The number comes from a startup launch post describing one of its clients, and the detail that stuck with me was the direction: the bill only goes up. That tracks, because what the bill buys is intelligence by the token, and every quarter a company finds new places to apply it. Intelligence became a utility. The question worth asking is the one nobody asks about electricity: which kind are you buying, and for what?

Where the tokens go

Watch where the intelligence actually flows in a business and a split shows up. Part of the spend creates: new code, new campaign drafts, new product specs, new proposals, work that didn't exist yesterday morning. The rest of it reviews: does this PR match our conventions, does this draft sound like us, does this clause fit our policy, is this ticket tagged right. Creation is the blank page. Review is the manager pass that decides whether the thing fits how we do it here.

The two kinds want different models.

Creation, especially the hard kind (an agent debugging a race condition, a deal memo on a company nobody has analyzed), rewards the absolute frontier. The task is novel, the ceiling is high, and every extra point of capability shows up in the output, so you rent the best model money can rent and you're right to.

Review runs the other way. The task is narrow and repeats a thousand times a month, and the bar is consistency with your standards, which live in your history, in every approve and reject your team has ever issued, inside no lab's training data. A frontier model reviewing your work is a brilliant stranger guessing at your taste. A model tuned on three months of your verdicts is a colleague who has read all of them.

Review has one more property that matters: it labels itself. Every artifact your team approves or sends back is a graded example, produced as a byproduct of doing the work. Creation data is expensive to make; judgment data piles up for free. Therefore the review side of your AI spend sits on exactly the asset that tuning needs.

The lab you can now rent

The reason companies haven't acted on this split is that owning a model used to mean becoming an AI lab: data pipelines, eval harnesses, training runs, ML hires who cost more than the savings. That wall came down in about twelve months. Thinking Machines took Tinker, its fine-tuning API, out of waitlist in December; Mira Murati's lab runs the GPUs, you bring the data. CoreWeave acquired OpenPipe last September to fold the capture-your-logs-and-tune flywheel into its cloud. The whole layer is moving from research capability to product you can sign up for.

The most complete version I've found is Understudy Labs, a YC company in the Summer 2026 batch, founded by two engineers who built experimentation systems at Instacart. The loop runs end to end: it captures your production traces (e.g. your agent logs), turns your successful runs into evals, post-trains small open-weight specialists on them, and routes traffic to a specialist only when it beats the model you're currently paying for on a held-out test. You own the resulting weights and can run them anywhere. Their published examples include a CRM agent scoring 13% higher than Sonnet 4.6 at a quarter of the cost, an 8B model matching it at 5.2x lower latency, and a sentiment pipeline at 50x lower cost. One line from the launch post deserves a wall: "Every company using AI should be accumulating intelligence, not accumulating API bills."

Faster, better, cheaper, conditionally

Usually there is a trade off between faster, better, and cheaper (the joke goes: pick-two), so it's worth being precise about when this one delivers all three. Cheaper is structural: small models cost less to run, full stop. Faster is structural too, and it reaches surprisingly high; Cursor trained its own Composer model with reinforcement learning for exactly one job, agentic coding inside its own editor, and got frontier-level results at four times the speed. Better is the conditional one. It arrives on one narrow task at a time, and only when your data is good: clean traces, consistent verdicts, enough volume to learn from. A review of 287 published case studies found tuned small models beating generic frontier ones on well-defined, domain-specific tasks at 10 to 100x lower running cost, and the operative words are well-defined. A specialist wins on the task it was tuned for, which happens to be the task your company runs ten thousand times a month, which happens to be where the money is.

If your traces are a mess, the eval gate is the honest part of the machine: the specialist simply never wins the bake-off, and you've learned something true about your data for the price of a failed experiment.

Readers with long memories will recognize a bill coming due. In July I wrote about owning your intelligence stack in layers, context and skills and evals all ownable, and I called the model itself "rented, for now." This is the for-now arriving. The Ramp Router that opened last week's post is the first rung of the same ladder, routing each request to the cheapest model that clears the bar. Tuning is the rung below it: make a cheaper model clear the bar, then own it.

The Monday test

Pull your AI spend and sort it by task instead of by vendor. Find the biggest line that repeats: the review pass, the classification, the tagging, the QA step that runs hundreds of times a week with a clear pass or fail. Then ask two questions about it. Are we keeping the traces and the verdicts? Would two more IQ points on this task change anything a tenth of the cost wouldn't? For the repeated lines, the honest answers are usually no and no.

Then start the cheapest version of the flywheel: keep every trace and every human verdict on that one task, starting this week, even if you tune nothing for a quarter. The verdicts are the training set. You're already paying to generate it, and the only question is whether you keep what you paid for.

The frontier labs will keep winning the work nobody has done before, and you should keep renting them for it. The judgment about how your company does things was never theirs to sell. Rent the pen. Own the judge.

Till next time,
Chris

Would you share with a friend?

Login or Subscribe to participate

Reply

Avatar

or to participate