Human in the Loop

THIS WEEK

Hi, it's Andreas here.
A lot happened in the last 14+ days, even by AI standards. Last winter OpenClaw, an open-source agent you run on your own machine, showed what a personal AI agent could be: always on, plugged into your apps. Now almost every big lab ships one: Google's Gemini Spark, Meta's Muse, SpaceXAI's Grok Bot, Microsoft's Copilot Autopilot, and last week OpenAI's dots, agents with their own cloud computer. Only Anthropic is still missing. My bet: that is what they ship next.

The third Agentic AI Cohort moved to next week so we can build in feedback from the last run (4 seats left).

In this issue:

  • The Briefing: Washington stops saying AI, OpenAI's safety report lead walks out, and ChatGPT opens up to agents

  • Thought Loop: how agents are rewiring software teams, and what I see in my own client work

  • Hands On: Andrew Ng's free prompting course, a GPT-6 model guide, open computer-use models and a tool to audit your own agents

THE BRIEFING

Trump launches a "Super Intelligence Force" and drops the word AI
→ Intelligence director Jay Clayton, now also AI czar, leads it. A Sep 29 executive order tells federal agencies to say "Super Intelligence" instead of "Artificial Intelligence."

OpenAI's safety report lead quits and calls the company's culture broken
→ David Robinson led the safety reports for OpenAI's major launches; he writes that trial-and-error deployment "guarantees periodic failures." Days earlier, OpenAI parted ways with three safety researchers.

OpenAI opens ChatGPT to third-party agents at DevDay
→ Plugins get their own panels inside ChatGPT, a new Agents API brings Codex's multi-agent tools into other apps, and OpenAI agents can now run inside AWS.

Anthropic commits $100M to train 10,000 engineers who deploy Claude inside companies
→ The Claude Frontier Academy pairs an in-person program with a 12-week residency on a real use case. Accenture, Bain, Deloitte and McKinsey are in the first cohort.

NVIDIA ships an agent safety stack: an open-source sandbox plus a hardware watchdog
→ OpenShell enforces policy on what an agent can touch; Sentry watches it from a separate chip. Anthropic, Microsoft, SAP and ServiceNow are named partners.

THOUGHT LOOP

How agents are rewiring software teams

For decades, programming meant translation: understand the problem in human terms, then render it in the syntax a machine can execute, down to the last curly brace and semicolon. Agents are fully removing that step. Autocomplete became suggestions, suggestions became chat, and now agents clone the repository, plan a change across files, run the tests and open the pull request without anyone typing a line. Developers say what they want built, not how to build it. The machine handles the complete implementation, and humans bring intent, architecture and judgment.

This is already happening in daily work. In JetBrains' 2026 survey of more than 15,000 professional developers, 90% use AI coding agents at work every week and 68% every day. They estimate that about 47% of their code is written entirely by agents. At Google, Sundar Pichai said in April that 75% of all new code is AI-generated and approved by engineers, up from 50% the previous fall.

A lot is being written about how this changes the software development life cycle (SDLC). Most of it stops at "the agent writes the code." But once the agent writes the code, where does the work go? Google and Anthropic, two companies that sell these tools, have measured it in their own teams.

The vendors measured themselves

Google's DORA team, which has benchmarked software delivery for a decade, surveyed almost 5,000 technology professionals for its 2025 report. 90% use AI at work, and more than 80% say it made them more productive. Teams shipped more, and broke more. It was worst where automated testing and fast feedback were weak. DORA's own summary: "AI doesn't fix a team; it amplifies what's already there."

Anthropic studied its own engineers. They now use Claude in 59% of their work, up from 28% a year earlier, and report a 50% productivity gain. But more than half say they can fully hand over only 0 to 20% of their work. One engineer put it this way: the job shifted "70%+ to being a code reviewer/reviser rather than a net-new code writer." And at the research frontier it looks the same. Claude leads 26% of Anthropic's AI research work, according to its newer R&D index. The share that runs fully on its own is still zero.

The speed gains are also smaller than the adoption suggests. METR, an independent AI evaluation lab, found experienced open-source developers 19% slower with AI in early 2025. Its 2026 follow-up shows a speedup, but the uncertainty range runs from 38% faster to 9% slower. The gain is not in typing speed. The work moved somewhere else.

What I see in my own work

Over the last two months, I helped several companies understand what all this fast, AI-written code does to their (software) development teams (and I also wrote in my upcoming book about it): how delivery changes, where forward deployed engineers fit, and whether they also need forward deployed units or strategists.

Before, most of the working day goes into writing code. Now the split point slides left and most of the day goes into judging it.

Same working day, different split. Proportions are illustrative, not measured.

First: the SDLC did not shrink. It moved. The classical SDLC stages are all still there: requirements, design, build, test, review, deploy, operate. Only build got faster, from weeks to hours, as Anthropic's agentic coding trends report puts it. Every stage around it got heavier: specs precise enough that an agent cannot misread them, reviews of code nobody on the team wrote, and tests that decide whether a change ships.

Before: work queues up in front of writing code. After: developer, spec, three agents, code, and the queue now sits in front of verification.

The queue used to form in front of writing code. Now several agents write in parallel from one spec, and it forms in front of verification.

Then there is everything the agent does that nobody asked for. This week Glow Labs reported that coding agents had pushed more than 13,000 internal screenshots from 300+ companies into public GitHub repos, just to attach images to pull requests. The code was probably fine, but nobody was checking what else the agent did.

For me, that is the line between vibe coding and (agentic) engineering: whether someone evaluated the code before it was deployed. In the companies I work with, almost nobody refers to their own AI-assisted coding as vibe coding. Vibe coding is always what the other team does (think about it for a second).

Second: the team itself changes shape. The classic team is a chain of specialists handing work to each other. The agentic team I see emerging is a small core: a product lead and a tech lead who decide what gets built and what ships, a few full-stack engineers who steer the agents, and agents doing the rest. We are very early here, and most companies have not started to think about that part, but I believe there is a major change coming very soon.

Classic team: 8 roles handing work down a chain. Agentic team: product lead, tech lead and two engineers, with agents doing code, tests, bugs, refactoring, docs, review and infra around them.

How I see teams changing: from a relay of eight specialists, each handing work to the next, to a core of four with agents doing the execution around them.

Where I think this is heading

The SDLC is changing because the bottleneck moved. Generating code is now cheap, and almost anyone with an agent can produce a working feature in an afternoon. What stays expensive is judgment and verification: deciding what is worth building, checking that it does what it should, and owning it once it runs. That is where people now spend most of their time.

The shape of teams will move with it. How exactly is hard to say today, and the core of four above is what I see now, not the final answer. But I am sure we will see some surprises: roles that shrink faster than anyone expects, roles that grow, and a few that do not have a name yet.

HANDS ON

Andrew Ng's AI Prompting for Everyone course on YouTube

Andrew Ng's new 7-hour prompting course is free on YouTube
→ 21 lessons on getting real results from ChatGPT, Claude and Gemini: when to use deep research, how to get honest pushback, and building apps without code.

OpenAI's guide to choosing between GPT-6 Astra, Sol and Luna
→ Free guide on which model fits which job, how to set reasoning effort, and how caching and compaction cut the cost of long-running work.

H Company opens Holo4, computer-use models you can run yourself
→ One model works across screens, code, MCP servers and APIs. Weights are on Hugging Face; the 35B is Apache 2.0, the 27B is non-commercial only.

Inspect Petri tests your own agents for out-of-scope behavior
→ An open-source auditing agent from Meridian Labs and the UK AI Security Institute's red team, with 170+ ready scenarios to run against your setup.

TWO WAYS I CAN HELP

I take on a small number of advisory engagements each quarter for AI programs that are stuck. Reply to this email with the one initiative you would want a second pair of eyes on. I read every reply.

And if you want to build agents rather than read about them, the Agentic AI Cohort starts its third run next week, with 4 seats left.

Andreas

Andreas Horn