This website uses cookies

Read our Privacy policy and Terms of use for more information.

Happy Tuesday ⚡️

Anthropic spent June asking Meta about renting computing power from Meta's data centers, up to $10 billion worth over two years. It may never sign, and that isn't really the point. Meta rents too. Its own filings commit it to buying up to $14.72 billion of cloud capacity from somebody else, on top of $237 billion in commitments that are mostly more of the same. Everyone at the top of this industry is renting from everyone else. So the "let's own our AI stack" pitch making the rounds inside your company deserves a much harder look.

Today, we're talking about:

  • A Chinese open-weight model beat Claude on a coding leaderboard, and almost everyone read the price tag wrong

  • Our ops lead got new filters onto our careers page from a Slack thread, and he doesn't write code

  • Plus one of the people who built Claude Code on why your best engineer going 10x doesn't move your company, and Mira Murati's lab giving its first model away

The Cheap Model Isn't Cheap

Our cofounder Arman went on Fox Business with Charles Payne Friday to talk about Kimi K3, with the leaderboard up on the screen behind him.

Moonshot AI's new model had just taken first place on Arena's Frontend Code board. That's an independent scoreboard where developers vote on which model writes better web code. K3 beat Claude Fable 5 and GPT-5.6 Sol, and it's the first Chinese model to hold the top spot. Moonshot has since stopped taking new subscriptions because too many people showed up.

Our engineers ran it all day before he went on air, and Alex wrote up what they found.

Everyone read the price tag wrong. K3 runs $3 per million tokens in, $15 out (a token is about three-quarters of a word). Against Fable 5 at $10 and $50, that's a rout.

Except nobody runs daily work on Fable. It's the most expensive model Anthropic sells. Line K3 up against Claude Sonnet 5, which is what teams actually use, and the price is identical: $3 and $15, to the dollar.

And nobody buys tokens. They buy finished jobs. Our engineers found K3 burns more tokens to finish the same job, which puts the total right back alongside the expensive American models.

Oh, and did we forget to mention that you can't run this thing?

The files land by July 27 and they're enormous even compressed. Moonshot's own guidance is 64 or more specialized chips. That's a rack of hardware, a line item on your budget, and somebody whose job is keeping it alive.

So all that free intelligence still runs on somebody else's servers, at a price they set.

Our take: this is real progress and it deserves the credit. An open-weight model pushing the frontier this hard drags the rest of the market along with it, on price and on roadmap, and that's good for everyone doing the buying.

(There are already whispers about Washington moving to slow the spread of Chinese open models. That's a story for another week.)

Still, open weights move the bill. They don't lower it. Off the lab's invoice, onto your infrastructure and your ops team. We're a long way from running all our own intelligence ourselves.

So if the pitch inside your building is "we'll switch to the open model and cut AI spend," ask what one finished job costs.

A year ago, owning your own intelligence looked like a someday problem. It's showing up faster than that.

Try it: K3 runs $3 in and $15 out per million tokens, through Moonshot or through OpenRouter, which resells a lot of models from one place. Weights land by July 27.

Sponsored

We rebuilt our entire inbound engine inside one Typeform.

Five different conversations hit our contact form every week — SMBs, enterprises, partnership pitches, workshop requests, candidates. Same generic form, same slow reply. We were losing opportunities to speed, not fit.

So we put all of it inside a single Typeform. Now it enriches every lead, figures out which bucket they belong in, and routes them to the right person in Slack — with a calendar link waiting on the success page for high-fit leads.

We described the logic in plain English; Typeform AI built the flow. Capture, enrich, route, close — one tool.

Multiplayer mode with Claude Tag

Brett runs ops here. He doesn't write code, and he's one of a handful of people on our team who need to make changes to the marketing website. Careers page, blog, marketing pages.

So we put all of them in one Slack channel with Claude Tag and one builder.

Any of them can ask Claude for a change now. The builder stays in the channel, keeps an eye on what ships, and gets people unstuck when Claude runs into trouble.

Brett asked for new filters on the careers page. Claude got stuck. JJ jumped into the thread and walked it through while Brett watched. The filters went live.

That part is new. Everyone in the channel can see what the agent did and correct it, so the work never disappears into somebody's private chat with a bot. Claude calls it multiplayer, and it’s awesome.

Two things make Claude Tag feel less like a chatbot.

Ambient mode means Claude stops waiting to be tagged. It flags what it thinks you need to know from the channels it's in.

Routines are standing work, and that's Anthropic's term for them. You name the schedule, like "every two hours, check the alerting dashboard and post anything new." It only posts when something changed.

Support email is the clearest use case. Skills can answer tickets today, but somebody has to go run the skill. With a routine, Claude sweeps the channel on the schedule you set, answers what it can, routes the rest, and tags a human on anything it shouldn't decide alone.

The setup steps are for whoever runs your Slack workspace, so forward them:

  1. The admin part is where everyone stalls. You need to be an Owner in your Claude organization. Start at claude.ai/admin-settings/claude-tag, add the app, then have a Slack workspace admin post @Claude connect as a brand new message, not a thread reply.

  2. Invite it into one channel first. Pick somewhere the work is real and a mistake costs you an apology instead of a customer.

  3. Write your standards down where it'll read them. Naming conventions and review checklists belong in your project's CLAUDE.md, which loads every session. The .claude/skills/ folder is for procedures it should reach for only when a task calls for one. Put permanent rules in a skill and they get skipped most of the time.

  4. Know what the guardrail actually is. Claude can only propose changes. It can't publish or launch anything on its own. But a proposal still kicks off your automated tests before a human reads a line, so somebody who understands those tests should review what runs on them.

  5. Treat channel membership as codebase access. Anyone who can post in that channel can steer the agent, and so can anything they paste into it. Keep it out of channels where customer emails and vendor messages land.

JJ recorded the whole setup, including the admin permissions that trip people up and the moment Claude got stuck on Brett's filters. It's free on YouTube. Steal it.

Your best engineer going 10x doesn't move your company — One of the people who built Claude Code says he hears the same thing everywhere. One person triples their output and the org doesn't budge. He lays out the five stages companies pass through on the way to fixing it, and argues most teams are tracking the wrong number entirely. Boris Cherny

Mira Murati's lab shipped its first model and gave the weights away — Thinking Machines released Inkling, which handles text, images, and audio, free for anyone to use or modify, and built to be customized to your specific work from day one. A frontier lab betting that fitting your job beats being generally smartest. Worth a look if you've been told your use case is too niche for a general model. Thinking Machines

DeepMind's CEO put a deadline on AI regulation — Demis Hassabis argues human-level AI is a few short years out and will land like ten Industrial Revolutions at ten times the speed. Then he proposes something concrete: a national AI safety watchdog modeled on the body that polices Wall Street, running before year-end, with labs handing over models 30 days before release for testing. Demis Hassabis

DoorDash built a version of itself with no app — It's for software to use instead of people, so an AI agent can search stores, compare deals, and check out with nothing on screen. A very normal company just decided its next customer might be a program, which is a question every consumer business is about to face. Andy Fang

What the people building AI agents actually argued about this year — The best field notes from this year's AI Engineer World's Fair, and the theme that keeps surfacing is that the expensive part has moved from the model to everything wrapped around it. Who reviews the work, where the rules live, who owns it when it breaks. One line worth the click, from Google DeepMind's Philipp Schmid: "Agents are just files." Latent Space

A free five-hour security course, for the people now shipping code — 83 lessons on finding weaknesses and defending against common attacks, aimed at working developers. Free right now. If AI is opening proposals in your codebase, somebody on your team should be able to read them suspiciously. Scrimba

Open roles:

  • AI Strategist

  • Forward Deployed Engineer

  • Applied AI Engineer

  • Engagement Manager

Salary ranges vary by role and experience. Additional comp based on output. Must be NY-based.

Keep Reading