
Happy Wednesday ⚡️
This week Chinese labs are flexing. Days after Dario called for a coordinated slowdown in AI while also arguing the US should impede China's progress, DeepSeek's CEO reportedly told investors it's planning a model five times the size of its current flagship. Alibaba announced its own massive model alongside a chip it says triples the performance of its predecessor. Xiaomi is pulling the curtain back, publicly training its next models in real-time with a meter showing its compute bill, performance, restarts and failures.
Speaking of transparency, this week Google finally confirmed that a Gemini model got into three companies' systems back in May, during a security test run by Irregular. An environment that was supposed to be offline connected to the internet. Oops. OpenAI published a framework for disclosing its own incidents, showing off that framework with six heretofore unreported examples. Evaluators from Accenture are en route to Anthropic, with each company pledging at least $1 billion over five years, but METR is in talks to pilot similar access with their own cash.
Meanwhile, a coalition of 20 countries led by Finland and Norway declared at UNGA that AI must stay under human direction and control. Seems reasonable. Of the 94% of the world whose governments didn't sign, notably absent were, you guessed it, the US and China. They got Iceland and Moldova on board though, so there's that. Trump announced an AI Force and will be hiring an "AI czar." Those interested should keep in mind that "Only High I.Q. individuals need apply." Hopefully the uniforms will make us look like the good guys.
Here's what else we're talking about around the Celsius fridge:
AI-assisted…garbage trucks?
How Chinese models can save you money
The first breach a regulator has logged with an AI agent as the attacker and more

What shipped this week and what it says about where this is going.
AI-assisted…garbage trucks?
Waymos aren't the only vehicles driving around with cameras on them. In Dallas, so are the garbage trucks. But instead of watching the road, they're watching houses, looking for code violations like illegal dumping and graffiti. The city mounted 100 cameras on 50 waste trucks, had AI software scan the photos, and turned its routes into a recurring citywide inspection.
Testing began in April, and by early September here are the numbers: 21,000 properties flagged, 3,700 courtesy notices sent, and 57 citations at 34 properties. Notices give owners a chance to fix a problem first, and no citation goes out until a code officer inspects the property in person. The trucks pass the same properties roughly every 30 days, so the city could measure whether flagged problems actually disappear. Instead, the only measurable target it set is 5,200 notices a yearafter full rollout, and it's already at 3,700.
The notices are a meaningful start. We don't know if it's deliberate, but there's something to be said for rolling out a program slowly. It probably isn't good politics to spray citations across the city from jump. That being said, not everyone is thrilled about the pilot, including one community that received a disproportionate number of citations.
Our take: Going slow is smart, but slow only pays off if you're learning something, and a notice count can't tell whether the program works. That's why we set a baseline for the outcome before anything gets automated: once a system is live, activity is the easiest number to report and the easiest to inflate. Measure what you actually want, or the system will get very good at producing whatever you counted.
Save your tokens
AI is particularly well-suited for small, repetitive jobs: tag a support ticket, pull a field from an invoice, decide whether a human needs to look at something. At volume, those jobs are where the per-token bill adds up, and they're where cheap open models, many of them Chinese, are often adequate.
In June, the founder of Lindy, a San Francisco AI-assistant startup, said he'd moved all of its traffic to DeepSeek and saved millions. His explanation: "You don't need God to write your email." He has since said the switch didn't hold for Lindy's more complex product, which tells you where the line is. Hosted on US infrastructure, open-weight models can also avoid sending your data to a Chinese company's API.
The model race keeps producing cheaper models that can handle more of your workload. They still won't handle your hardest reasoning or your longest agent workflows, but isolate the high-volume steps, and a cheaper model can cut your costs and give your closed vendor less room at renewal.

See what kind $$$ Chinese models can save you
Pick one high-volume workflow and randomly pull 100 requests from last week. Define one concrete pass/fail test, ideally the bar you already accept in production. Then run all 100 through your current model and one cheap open model on US-hosted infrastructure, and count the passes for each.
The number that matters is the gap. If your current model passes 92 and the cheap one passes 89 at a fifth of the cost, you're paying five times as much for three points. Multiply by your real volume and pricing, then decide whether those three points matter for this workflow.
If they don't, move it. If some failures are easy to catch automatically, like malformed output, send just those back to your current model. Either way, you walk into your next renewal with a number.
Twelve hundred agents organized themselves in five days. Most companies can't get two agents to hand off a task without a person in the middle.
Tenex is the AI engineering team behind this awesome newsletter. We build the boring layer that makes agents safe: identity, permissions, and a record the agent didn't write itself. About to put agents somewhere that matters? Let's talk.

In case you missed it
Here's why Jev blew up this week: it became the fastest-adopted model in Vercel's gateway history without writing a word. TypeSafe AI's model returns typed decisions instead of prose (a choice, a score, or a yes/no, each with a probability attached) and no reasoning text. It reached nearly 13% of paid teams within a day, more than six times Fable 5.1.
Anthropic says Claude now leads 26% of its AI R&D work, up from under 1% in February, and is at least a collaborator on more than 90% of it. None of it is fully autonomous, and this is Anthropic measuring its own R&D, so treat it accordingly, but still impressive.
OpenAI shipped Astra for Law, a legal configuration of GPT-6 Astra, at 54% overall correctness on its own test versus 38.7% for the standard model with web search. The interesting part is that they published the ceiling.
Spain's data regulator says it received its first breach notification involving an AI agent carrying out an attack. The agent scanned files, logged in, found a vulnerability, modified personal data and read invoices. Separately, GreyNoise saysan attacker used hundreds of AI agents to exploit two PaperCut flaws across 395 organizations in 48 countries. Meanwhile, a Cisco survey found 85% of organizations are experimenting with or piloting agentic AI, while just 5% have broad production deployments.
You can tell your AI coding tool to use only the exact version of an add-on you approved. It turns out none of the major tools actually checked. Someone else's code could be installed and run without you clicking anything. Claude Code and Codex are patched. Microsoft hasn't fixed Copilot, and Google told the researchers it won't patch Gemini CLI because it's deprecating it anyway.
Stanford researchers built a pipeline that turns a scientific paper into software an AI agent can actually call. It converted 74 of 100 computational biology papers with no human involvement; in one case study, building an agent took about 45 minutes and $14 of compute. Three of these paper-agents, chained together, pointed to a gene called GPR137 as the probable culprit behind a psoriasis-linked variant, and the result matched existing experimental data once a researcher chose how to check it.
AT&T and Crocs said at Dreamforce that agentic AI has been oversold and that the value lands in narrow, well-scoped cases. Digital Realty surveyed 2,131 IT decision-makers and found that 40% now name a lack of specialized infrastructure as their main constraint, up from 9% in 2024.
Entroducing just turned 30. Try not to let anyone catch you zoning out while you listen.
See you next week!

Open roles:
AI Strategist
Forward Deployed Engineer
Applied AI Engineer
Recruiter
Salary ranges vary by role and experience. Additional comp based on output. Must be NY-based.


