This website uses cookies

Read our Privacy policy and Terms of use for more information.

Happy Wednesday ⚡️

On Saturday, Dario said (again) that the industry needs to slow down so safety work can catch up. What's unusual was the reaction: within hours, industry rivals who don't exactly brunch together co-signed. Washington and Beijing chimed in “no chance” almost in unison. The US wants to stay ahead of China, and China wants, well, the opposite of that, so a coordinated slowdown seems…unlikely. 

Global cooperation or not, Dario committed to open Anthropic’s systems to outside evaluators who can publish findings without approval. Sam says OpenAI will do the same and Elon agreed in the abstract. Which systems, which evaluators, or when remains anybody’s guess. 

Concurrently, another AI researcher pulled the fire alarm. This time it’s Jacob Coxon suggesting there is a non-zero chance that AI will kill us all. Hopefully we’ll be able to predict the fate of the world in the next Ultrathink. Until then, here’s what we’re talking about:

  • Another AI giant is deploying an army of engineers to get their customers’ AI systems up and running effectively

  • The TSA is deploying extremely effective AI agents (that don’t tell you to take off your shoes)

  • How to tell which of your workflows an agent could actually do, and how to write down the rules for one.

  • How Grok can help you get your errands done on time, for once.

What shipped this week and what it says about where this is going.

How many engineers does it take… 

Another major player is putting boots on the ground. Google Cloud and Accenture announced last week that they're standing up a joint business group with up to a thousand forward-deployed engineers, whose job is getting Gemini Enterprise running inside client operations. This follows AWS, Microsoft, Anthropic and OpenAI who all now see the need for real hands-on help. It’s worth mentioning that if orgs just needed more powerful models, none of this would be necessary. 

You can keep your shoes on

You probably wouldn’t have imagined the TSA would be deploying successful AI agents before much of the business world, but here we are. Their AI agent Ace, built on Salesforce Public Sector Solutions, is now handling roughly 100,000 traveler conversations a month, resolving 96% of routine inquiries without a human. What percent of all inquiries are routine is an open question, but the bigger find is that interaction costs are down more than 90% from above four dollars a conversation.  

Ace is built on prior TSA projects that pulled traveler questions from email, phone and social media into one system, and it was built with the agency, on a platform, for one narrowly defined job. It works because the answers exist in writing and there is little ambiguity. 

Salesforce is betting it can pre-build that kind of agent for the most common jobs. Last week it shipped seven named agents built for specific roles across service, commerce, HR, supply chain, and starting November, sales. Though pre-built, the agents still have to be configured to the customer's data and policies before it does anything.

Our take: The big jumps in performance and deployability come from systems built for specific roles. A purpose-built system shows up knowing the role's inputs, its rules, the tools it can touch, and what a good outcome looks like. A general agent has to be told all of that, or guess. And once you can say what good looks like, you can test for it and fix what misses.

Salesforce's answer is to sell you the role pre-built. Ours is the bet we've been making from jump: custom systems, built and deployed by the same team, so the people who learned how the work gets done are the ones who encode it. Ace, for what it's worth, is closer to that than to anything in a box.

Automate what’s boring 

For now, forget what AI might be able to do and focus on what it can do well enough to justify the spend today. TSA’s Ace agent works because every question put to it has a documented, correct answer. Here’s how you might pull your own Ace off. 

  1. Consider the simplest workflows that come up for you or your team and ask: Could someone write down what an objectively correct outcome looks like, and precisely enough that an outsider could replicate the work? 

    • Yes: this is a good workflow to try automating. 

    • No: find a simpler, more objective aspect of a workflow to use. 

  2. Pull completed cases where you know the outcome was right, and pull the inputs too, including several outliers. You need what the person could see when they decided, not the file after everything resolved. 

  3. Prepare the cases in a file folder (or single document) and upload them to an assistant like Claude Co-Work or ChatGPT Work. Input this prompt or some version of it to fit your project:

    • Attached are completed cases from the same workflow. Each includes the information that was available when the decision was made, and the final outcome. Infer the decision rules that were actually applied. Write them as explicit if-then statements. Flag every case where two similar inputs produced different outcomes, and state what additional information would have been needed to predict which. List the rules you're least confident about and explain why.

  4. Mark every rule that's wrong. The corrections are where the undocumented exceptions surface. The corrected list is your spec and your test set.

Figuring out how the work actually gets done, which rules people follow without knowing they're rules, is where a forward-deployed engineer starts. This is a version of that you can run yourself on a workflow you already own.

Twelve hundred agents organized themselves in five days. Most companies can't get two agents to hand off a task without a person in the middle.

Tenex is the AI engineering team behind this awesome newsletter. We build the boring layer that makes agents safe: identity, permissions, and a record the agent didn't write itself. About to put agents somewhere that matters? Let's talk.

🧬 DeepMind released AlphaGenome Atlas, which predicts what every possible single-letter DNA change does. Broad Institute researchers used it on unsolved rare-disease cases and it caught one earlier work had missed, in a gene tied to childhood epilepsy. Lab tests confirmed it.

🙅 Twenty-five Fields Medalists told AI labs to stop treating problem-solving as the point. Their September 11 declaration says solving a problem is only a proxy for understanding it, and that results dumped out at speed erase the writeup and credit that make the field cumulative. It landed three days after OpenAI said 10,000 agents running 88 hours had cracked a version of a famous open problem.

😬 Anthropic's threat report found a weapons cell in northern Yemen using Claude Code instead of software engineers to write rocket guidance code. They test-fired one, it failed, and they came back hours later to debug. Six weapons cases in all, across Yemen, Russia and China. Anthropic says it found no evidence anyone fielded a working weapon.

🛠️ A project folder someone hands you can run code on your laptop the moment you open it with an AI coding tool. No prompt, no warning, and the AI doesn't have to do anything wrong; the tool's own background git commands run outside the sandbox, and git executes whatever the repo's config says. Claude Code and Cursor CLI were both affected and both are patched. Run claude --version and make sure you're on 2.1.247 or later. Full writeup.

⛑️ Also from last week: two more safety researchers left Anthropic and Google DeepMind for METR, arguing a lab could lose control of a system and nothing would require it to say so. California signed an AI auditor registry. Hawley opened an investigation into OpenAI's breach response, answers due October 1. And Paul Christiano joined OpenAI's safety committee having put roughly 4% odds on catastrophic loss of control within a year – about the odds of a white Christmas in Washington DC. 

Here’s a song to help calm your nerves, should you need it. 

This week’s edition of Ultrathink has been brought to you by Matthew Siegel, your new Lead Content Engineer at Tenex. I’m excited to be keeping you up to speed with what matters most in AI.

See you next Tuesday!

Open roles:

  • AI Strategist

  • Forward Deployed Engineer

  • Applied AI Engineer

  • Recruiter

Salary ranges vary by role and experience. Additional comp based on output. Must be NY-based.